Publications of NIPGR Scientists

Permanent URI for this communityhttps://ndkr-library.nipgr.ac.in/handle/123456789/1

Browse

Search Results

Now showing 1 - 10 of 27
  • Item
    AtFusionDB: A comprehensive database of fusion transcripts in model plant Arabidopsis thaliana
    (Springer Nature Publishing AG, 2026) Shree, Tanu; Kumar, Shailesh
    Fusion transcripts are chimeric RNAs, produced by the joining of two different RNAs at the RNA level or as a product of gene fusion at the DNA level. In this era of high-throughput sequencing technologies, it is easy to identify novel molecules like fusion transcripts in different systems. That's because, initially, supposed to be the well-known cancer biomarkers, fusion transcripts are also validated in normal human physiology. In Planta, discrete reports are available, indicating the presence of fusion transcripts but no dedicated web resource is available for the plant-specific fusion transcripts. This chapter describes the first plant-specific database of fusion transcripts, i.e., AtFusionDB ( http://www.nipgr.res.in/AtFusionDB ), which contains the information on fusion transcripts identified in the model plant Arabidopsis thaliana. This database can be exploited to get significant information about gene/transcript fusion in plants.
  • Item
    Validation of plant fusion peptides using proteomics data
    (Springer Nature Publishing AG, 2026) Hamid, Fiza; Aftab, Sahrish; Shree, Tanu; Kumar, Shailesh
    Fusion transcripts and their fused protein products are emerging as exciting entities in molecular biology, offering potential applications in diagnostics and therapeutics. These fusion proteins, derived from the translation of fusion transcripts, hold promise as unique biomarkers and targets for intervention. While numerous algorithms exist to identify fusion RNAs, the detection and validation of their protein counterparts through proteomics remains a growing area of research. This challenge is particularly intriguing in plant biology, where fusion events may affect stress responses, development, and adaptation. This chapter provides an accessible and practical workflow for validating plant fusion peptides using publicly available proteomics datasets.
  • Thumbnail Image
    Item
    Uncovering the biosynthetic potential of Amycolatopsis: new insights into glycopeptide antibiotic and polyketide gene clusters
    (Oxford University Press, 2026) Bisht, Niyati; Mayilraj, Shanmugam; Kaur, Navjot; Kumar, Shailesh
    Background: : Amycolatopsis species are renowned producers of a vast array of biologically active molecules, including Glycopeptide antibiotics (GPAs), polyketides, siderophores, and terpenes. Despite their clinical significance, the full biosynthetic genetic capacity and evolutionary diversification of Amycolatopsis remain unexplored. Methods and Results: We analyzed 16 Amycolatopsis strains, including six newly sequenced in this work, six from our previously published datasets, and four retrieved from NCBI. Phylogenetic, pangenome, and antiSMASH-based genome-mining analyses were performed to identify secondary metabolite gene clusters, with a focus on NRPS, PKS, terpenes, and siderophores. Conserved glycopeptide gene clusters found across Cluster A strains, encoding core NRPSs, P450 oxygenases, and tailoring enzymes with variations consistent with the structural GPA types. Analysis showed conserved but distinct GPA BGC organization corresponding to the type I, II, and III subclasses, as well as their genetic, structural, and functional diversifications. A. azurea DSM 43854T produced A35512B rather than azureomycins, while A. alba DSM 44262T produced vancomycin. Six previously unreported Cluster A strains were found to encode putative GPA gene clusters, and LC–MS profiling predicted GPA production of nogabecin from A. keratiniphila subsp. keratiniphila DSM 44409T and A33512B from A. thailandensis JCM 16380T. GPA biosynthetic capacity was largely restricted to Cluster A, but in Cluster C, in the case of A. balhimycina DSM 44591T. Type II PKS, siderophore, and terpene gene clusters were also explored for these strains. Conclusions: This study provides a comparative genomic overview of Amycolatopsis Cluster A, highlighting GPA diversity and revealing broader potential for secondary metabolites.
  • Thumbnail Image
    Item
    PFGPred: A stack ensemble classifier for the identification of fusion genes in plants
    (Oxford University Press, 2026) Hamid, Fiza; Mukherjee, Kanka; Chaudhary, Sakshi; Kaushik, Love; Kumar, Shailesh
    Fusion genes play crucial roles in plant biological processes but remain far less explored than their human counterparts, largely due to limited validated datasets and the absence of plant-specific prediction tools. Existing approaches often produce high false-positive rates, restricting reliable discovery. To address this gap, we developed Plant Fusion Gene Predictor (PFGPred), an ensemble machine learning framework that integrates Random Forest, XGBoost, and long short-term memory (LSTM) models into a meta-classifier for accurate identification of true and false fusion genes from RNA-Seq data. PFGPred was trained on a high-confidence dataset of fusion genes validated by both RNA-Seq and whole-genome sequencing from Arabidopsis thaliana, Oryza sativa, Triticum aestivum, and Zea mays, to predict and rank candidate fusion genes for future functional validation. It outperformed individual baseline models, achieving accuracies of 0.97 on training data and 0.77 on independent test data. When evaluated on human datasets, it achieved 0.71 accuracy with lower sensitivity, reflecting biological differences between plant and human fusion events. Comparative analyses confirmed that PFGPred reliably identifies validated fusions, demonstrating its utility as a cost-effective, plant-specific prediction tool for high-throughput fusion gene screening and functional genomics research. It is freely available as a web server at http://www.nipgr.ac.in/PFGPred.
  • Thumbnail Image
    Item
    UBA1-CDK16: A female-specific chimeric RNA emerging through evolution and involved in immune regulation
    (American Association for the Advancement of Science, 2026) Shi, Xinrui; Blackburn, Loryn; Singh, Sandeep; Glowczyk-Gluc, Martyna; Tajammal, Anam; Zahra, Shafaque; Kumar, Shailesh; Cornelison, Robert; Liang, Chen; Qin, Fujun; Liu, Aiqun; Lin, Shitong; Tang, Yue; Elfman, Justin; Manley, Thomas; Bullock, Timothy; Haverstick, Doris M.; Wu, Peng; Li, Hui
    Chimeric RNAs resulting from intergenic splicing represent a distinct mechanism for transcriptome expansion. To explore the role of this previously unidentified layer of the transcriptome in sex-specific immunity, we analyzed RNA sequencing data from 425 blood samples and identified a female-specific chimeric RNA, UBA1-CDK16, which was further validated in more than 1200 blood samples. This chimeric RNA forms via cis-splicing between two adjacent X-linked parental genes, UBA1 and CDK16, despite both being expressed in both sexes. We demonstrated that a female-specific chromatin loop at the UBA1-CDK16 junction sites facilitates the intergenic splicing. Evolutionary analysis revealed that UBA1-CDK16 became female specific in humans through at least two independent paths. Functional studies suggested that UBA1-CDK16 is enriched in the myeloid lineage and may regulate myeloid cell development. Notably, its abnormal expression in female patients with COVID-19 correlates with altered neutrophil counts, highlighting its potential role in the disease progression.
  • Thumbnail Image
    Item
    AquaaG: A comprehensive pipeline for quality assessment and annotation of genomes
    (Elsevier B.V., 2026) Shukla, Jagriti; Mukherjee, Kanka; Sahu, Namrata; Kumar, Shailesh
    The rapid expansion of publicly available genome assemblies has made genome annotation an increasingly challenging task, particularly for large-scale analyses across prokaryotic and eukaryotic organisms. While several tools exist for assembly evaluation and annotation, their use often involves fragmented workflows that require extensive manual coordination. To overcome this limitation, we introduce AquaaG, an automated and reproducible genome annotation pipeline. AquaaG integrates genome assembly retrieval from NCBI, assembly quality assessment using QUAST, organism-specific annotation using Prokka for prokaryotes and BRAKER3 for eukaryotes, gene-space completeness evaluation using BUSCO, and functional annotation using EggNOG-mapper. The pipeline is configured through simple YAML files and supports species-level, kingdom-level, and custom assembly-based analyses with optional submitter-based filtering. AquaaG therefore provides a practical and reproducible framework for high-throughput genome annotation and assessment.
  • Item
    The intersection of AI and genomics in health and disease: Advancements and applications
    (Elsevier B.V., 2026) Kaushik, Love; Vivek, A T; Arora, Simran; Hamid, Fiza; Mukherjee, Kanka; Bisht, Niyati; Chaudhary, Sakshi; Shukla, Jagriti; Nawani, Sakshi; Kumar, Shailesh
    AI and genomics are revolutionizing precision medicine by using machine learning (ML) to analyze large-scale next-generation sequencing (NGS) data, identifying genetic mutations and biomarkers for personalized therapies. In practice, this accelerates drug discovery and enhances variant detection, while in cancer genomics, AI enables early detection via liquid biopsies and refines treatment by integrating multi-omics data to improve therapeutic precision. However, challenges such as data biases in underrepresented populations, limited model interpretability, and ethical concerns regarding privacy and algorithmic inequity hinder clinical adoption and demand robust governance. Efforts to diversify datasets also face standardization hurdles, although explainable AI and federated learning provide promising solutions for improving transparency and privacy. In this chapter, we discuss the role of AI in advancing genomics from diagnostics to novel therapies and emphasize the need for equitable frameworks to ensure responsible implementation, thereby paving the way for breakthroughs in personalized medicine.
  • Thumbnail Image
    Item
    AraNSdb: a dedicated database of stress-responsive non-coding RNAs in Arabidopsis thaliana
    (Springer Nature Publishing AG, 2026) Vivek, A.T.; Bhatia, Manika; Sahu, Namrata; Kalakoti, Garima; Kaushik, Love; Mukherjee, Kanka; Kumar, Shailesh
    Plants, as sessile organisms, are constantly exposed to biotic and abiotic stresses, making their ability to respond crucial for survival. Non-coding RNAs (ncRNAs) have emerged as key regulators in these stress responses, with several studies identifying numerous stress-responsive ncRNAs (SRNs). However, a comprehensive collection of SRNs derived from sequencing data in Arabidopsis thaliana has been lacking. To address this, we utilized high-throughput experimental data and mined published studies to construct AraNSdb (Arabidopsis ncRNA Stress Database), a systematic resource for storing and querying SRNs. AraNSdb documents over 1,000 expression profiles from diverse stress datasets, encompassing 6,616 SRNs, including microRNAs (miRNAs), small interfering RNAs (siRNAs), long non-coding RNAs (lncRNAs), and circular RNAs (circRNAs). The database features an intuitive web interface for exploring SRNs associated with specific stress types and provides detailed ncRNA annotations to support functional and regulatory studies. AraNSdb offers a valuable platform for advancing our understanding of ncRNA-mediated stress responses and is freely accessible at http://www.nipgr.ac.in/AraNSdb.
  • Thumbnail Image
    Item
    Integrative multi-omics analysis widens annotation and functional insights into long non-coding RNAs of Arabidopsis thaliana
    (Springer Nature Publishing AG, 2026) Vivek, AT; Kiran, Harikumar; Sahu, Namrata; Kalakoti, Garima; Kumar, Shailesh
    Background:- Long non-coding RNAs (lncRNAs) play key roles in regulating plant growth, development, and stress responses. Despite their increasing identification in plant transcriptomes, a systematic characterization of lncRNAs is still lacking, leaving a significant knowledge gap. To address this, we systematically identified and characterized Arabidopsis lncRNAs through integrative analysis of strand-specific RNA sequencing data and multi-omics datasets, revealing their genomic features, regulatory interactions, and evolutionary characteristics. Results:- Using a custom pipeline applied to hundreds of stranded RNA-seq datasets, we assembled a comprehensive catalog of 4,772 intergenic and antisense Arabidopsis lncRNAs. In comparing multiple key features of lncRNAs with those of protein-coding genes, we found that intergenic lncRNAs contain high transposable element-derived fragments and display broader TE diversity. Distinct DNA methylation and histone modification signatures further distinguished lncRNAs from protein-coding genes. We additionally uncovered R-loop connections and associations with sRNAs involved in post-transcriptional regulation and RNA-directed DNA methylation, with a minor subset classified as Pol V–transcribed. Of note, our results revealed lncRNAs mediating stress-responsive cis interactions and others linked to trait-associated loci. Probing further, an experimental evidence resource confirmed small peptide production from multiple lncRNA loci. Extending our investigation, comparative analyses across Brassicaceae species revealed syntenic lncRNAs enriched for shared sequence motifs despite substantial sequence divergence. Conclusions:- This study provides a valuable and extensively annotated catalog of Arabidopsis lncRNAs, revealing their diverse genomic features, regulatory interactions, and evolutionary characteristics. Altogether, our work advocates for multi-omics integrative analysis as a potent strategy to efficiently enhance lncRNA annotation, providing insights into functionality and addressing annotation limitations. Our comprehensive bioinformatic analyses of Arabidopsis lncRNAs pave the way for future functional characterization of these transcripts.
  • Thumbnail Image
    Item
    Breaking and making genes: the genesis of novel traits in plants
    (John Wiley & Sons, 2026) Hamid, Fiza; Arora, Simran; Kumar, Shailesh
    Understanding the mechanisms by which plants adapt, evolve, and acquire new traits is crucial for enhancing agricultural resilience and productivity in the face of global challenges. Among the various mechanisms that drive new gene evolution, gene fusion has emerged as a significant yet relatively understudied contributor. It can arise through chromosomal rearrangements or RNA processing mechanisms, merging segments from different genes to produce novel fusion transcripts. In plants, these fusion events have been associated with key biological functions, including the regulation of specialized metabolism, stress responses, and developmental changes. While fusion genes have been extensively studied in humans, mainly due to their oncogenic potential, their prevalence and functional relevance in plants remain relatively underexplored. This review offers a detailed overview of the molecular mechanisms underlying gene fusion formation, highlighting their participation in gene evolution, functional diversification, and plant adaptation. In addition, we discuss current methodologies for detecting and validating fusion events, including high-throughput sequencing technologies and emerging single-cell sequencing platforms, and outline promising directions for future research aimed at elucidating their biological significance. Collectively, these insights emphasize the expanding importance of gene fusions in plant biology and underscore the need for further investigation into their regulatory and evolutionary roles.