Institutional Publications
Permanent URI for this collectionhttps://ndkr-library.nipgr.ac.in/handle/123456789/11
Browse
48 results
Search Results
Item Uncovering the biosynthetic potential of Amycolatopsis: new insights into glycopeptide antibiotic and polyketide gene clusters(Oxford University Press, 2026) Bisht, Niyati; Mayilraj, Shanmugam; Kaur, Navjot; Kumar, ShaileshBackground: : Amycolatopsis species are renowned producers of a vast array of biologically active molecules, including Glycopeptide antibiotics (GPAs), polyketides, siderophores, and terpenes. Despite their clinical significance, the full biosynthetic genetic capacity and evolutionary diversification of Amycolatopsis remain unexplored. Methods and Results: We analyzed 16 Amycolatopsis strains, including six newly sequenced in this work, six from our previously published datasets, and four retrieved from NCBI. Phylogenetic, pangenome, and antiSMASH-based genome-mining analyses were performed to identify secondary metabolite gene clusters, with a focus on NRPS, PKS, terpenes, and siderophores. Conserved glycopeptide gene clusters found across Cluster A strains, encoding core NRPSs, P450 oxygenases, and tailoring enzymes with variations consistent with the structural GPA types. Analysis showed conserved but distinct GPA BGC organization corresponding to the type I, II, and III subclasses, as well as their genetic, structural, and functional diversifications. A. azurea DSM 43854T produced A35512B rather than azureomycins, while A. alba DSM 44262T produced vancomycin. Six previously unreported Cluster A strains were found to encode putative GPA gene clusters, and LC–MS profiling predicted GPA production of nogabecin from A. keratiniphila subsp. keratiniphila DSM 44409T and A33512B from A. thailandensis JCM 16380T. GPA biosynthetic capacity was largely restricted to Cluster A, but in Cluster C, in the case of A. balhimycina DSM 44591T. Type II PKS, siderophore, and terpene gene clusters were also explored for these strains. Conclusions: This study provides a comparative genomic overview of Amycolatopsis Cluster A, highlighting GPA diversity and revealing broader potential for secondary metabolites.Item PFGPred: A stack ensemble classifier for the identification of fusion genes in plants(Oxford University Press, 2026) Hamid, Fiza; Mukherjee, Kanka; Chaudhary, Sakshi; Kaushik, Love; Kumar, ShaileshFusion genes play crucial roles in plant biological processes but remain far less explored than their human counterparts, largely due to limited validated datasets and the absence of plant-specific prediction tools. Existing approaches often produce high false-positive rates, restricting reliable discovery. To address this gap, we developed Plant Fusion Gene Predictor (PFGPred), an ensemble machine learning framework that integrates Random Forest, XGBoost, and long short-term memory (LSTM) models into a meta-classifier for accurate identification of true and false fusion genes from RNA-Seq data. PFGPred was trained on a high-confidence dataset of fusion genes validated by both RNA-Seq and whole-genome sequencing from Arabidopsis thaliana, Oryza sativa, Triticum aestivum, and Zea mays, to predict and rank candidate fusion genes for future functional validation. It outperformed individual baseline models, achieving accuracies of 0.97 on training data and 0.77 on independent test data. When evaluated on human datasets, it achieved 0.71 accuracy with lower sensitivity, reflecting biological differences between plant and human fusion events. Comparative analyses confirmed that PFGPred reliably identifies validated fusions, demonstrating its utility as a cost-effective, plant-specific prediction tool for high-throughput fusion gene screening and functional genomics research. It is freely available as a web server at http://www.nipgr.ac.in/PFGPred.Item UBA1-CDK16: A female-specific chimeric RNA emerging through evolution and involved in immune regulation(American Association for the Advancement of Science, 2026) Shi, Xinrui; Blackburn, Loryn; Singh, Sandeep; Glowczyk-Gluc, Martyna; Tajammal, Anam; Zahra, Shafaque; Kumar, Shailesh; Cornelison, Robert; Liang, Chen; Qin, Fujun; Liu, Aiqun; Lin, Shitong; Tang, Yue; Elfman, Justin; Manley, Thomas; Bullock, Timothy; Haverstick, Doris M.; Wu, Peng; Li, HuiChimeric RNAs resulting from intergenic splicing represent a distinct mechanism for transcriptome expansion. To explore the role of this previously unidentified layer of the transcriptome in sex-specific immunity, we analyzed RNA sequencing data from 425 blood samples and identified a female-specific chimeric RNA, UBA1-CDK16, which was further validated in more than 1200 blood samples. This chimeric RNA forms via cis-splicing between two adjacent X-linked parental genes, UBA1 and CDK16, despite both being expressed in both sexes. We demonstrated that a female-specific chromatin loop at the UBA1-CDK16 junction sites facilitates the intergenic splicing. Evolutionary analysis revealed that UBA1-CDK16 became female specific in humans through at least two independent paths. Functional studies suggested that UBA1-CDK16 is enriched in the myeloid lineage and may regulate myeloid cell development. Notably, its abnormal expression in female patients with COVID-19 correlates with altered neutrophil counts, highlighting its potential role in the disease progression.Item AquaaG: A comprehensive pipeline for quality assessment and annotation of genomes(Elsevier B.V., 2026) Shukla, Jagriti; Mukherjee, Kanka; Sahu, Namrata; Kumar, ShaileshThe rapid expansion of publicly available genome assemblies has made genome annotation an increasingly challenging task, particularly for large-scale analyses across prokaryotic and eukaryotic organisms. While several tools exist for assembly evaluation and annotation, their use often involves fragmented workflows that require extensive manual coordination. To overcome this limitation, we introduce AquaaG, an automated and reproducible genome annotation pipeline. AquaaG integrates genome assembly retrieval from NCBI, assembly quality assessment using QUAST, organism-specific annotation using Prokka for prokaryotes and BRAKER3 for eukaryotes, gene-space completeness evaluation using BUSCO, and functional annotation using EggNOG-mapper. The pipeline is configured through simple YAML files and supports species-level, kingdom-level, and custom assembly-based analyses with optional submitter-based filtering. AquaaG therefore provides a practical and reproducible framework for high-throughput genome annotation and assessment.Item AraNSdb: a dedicated database of stress-responsive non-coding RNAs in Arabidopsis thaliana(Springer Nature Publishing AG, 2026) Vivek, A.T.; Bhatia, Manika; Sahu, Namrata; Kalakoti, Garima; Kaushik, Love; Mukherjee, Kanka; Kumar, ShaileshPlants, as sessile organisms, are constantly exposed to biotic and abiotic stresses, making their ability to respond crucial for survival. Non-coding RNAs (ncRNAs) have emerged as key regulators in these stress responses, with several studies identifying numerous stress-responsive ncRNAs (SRNs). However, a comprehensive collection of SRNs derived from sequencing data in Arabidopsis thaliana has been lacking. To address this, we utilized high-throughput experimental data and mined published studies to construct AraNSdb (Arabidopsis ncRNA Stress Database), a systematic resource for storing and querying SRNs. AraNSdb documents over 1,000 expression profiles from diverse stress datasets, encompassing 6,616 SRNs, including microRNAs (miRNAs), small interfering RNAs (siRNAs), long non-coding RNAs (lncRNAs), and circular RNAs (circRNAs). The database features an intuitive web interface for exploring SRNs associated with specific stress types and provides detailed ncRNA annotations to support functional and regulatory studies. AraNSdb offers a valuable platform for advancing our understanding of ncRNA-mediated stress responses and is freely accessible at http://www.nipgr.ac.in/AraNSdb.Item Integrative multi-omics analysis widens annotation and functional insights into long non-coding RNAs of Arabidopsis thaliana(Springer Nature Publishing AG, 2026) Vivek, AT; Kiran, Harikumar; Sahu, Namrata; Kalakoti, Garima; Kumar, ShaileshBackground:- Long non-coding RNAs (lncRNAs) play key roles in regulating plant growth, development, and stress responses. Despite their increasing identification in plant transcriptomes, a systematic characterization of lncRNAs is still lacking, leaving a significant knowledge gap. To address this, we systematically identified and characterized Arabidopsis lncRNAs through integrative analysis of strand-specific RNA sequencing data and multi-omics datasets, revealing their genomic features, regulatory interactions, and evolutionary characteristics. Results:- Using a custom pipeline applied to hundreds of stranded RNA-seq datasets, we assembled a comprehensive catalog of 4,772 intergenic and antisense Arabidopsis lncRNAs. In comparing multiple key features of lncRNAs with those of protein-coding genes, we found that intergenic lncRNAs contain high transposable element-derived fragments and display broader TE diversity. Distinct DNA methylation and histone modification signatures further distinguished lncRNAs from protein-coding genes. We additionally uncovered R-loop connections and associations with sRNAs involved in post-transcriptional regulation and RNA-directed DNA methylation, with a minor subset classified as Pol V–transcribed. Of note, our results revealed lncRNAs mediating stress-responsive cis interactions and others linked to trait-associated loci. Probing further, an experimental evidence resource confirmed small peptide production from multiple lncRNA loci. Extending our investigation, comparative analyses across Brassicaceae species revealed syntenic lncRNAs enriched for shared sequence motifs despite substantial sequence divergence. Conclusions:- This study provides a valuable and extensively annotated catalog of Arabidopsis lncRNAs, revealing their diverse genomic features, regulatory interactions, and evolutionary characteristics. Altogether, our work advocates for multi-omics integrative analysis as a potent strategy to efficiently enhance lncRNA annotation, providing insights into functionality and addressing annotation limitations. Our comprehensive bioinformatic analyses of Arabidopsis lncRNAs pave the way for future functional characterization of these transcripts.Item Breaking and making genes: the genesis of novel traits in plants(John Wiley & Sons, 2026) Hamid, Fiza; Arora, Simran; Kumar, ShaileshUnderstanding the mechanisms by which plants adapt, evolve, and acquire new traits is crucial for enhancing agricultural resilience and productivity in the face of global challenges. Among the various mechanisms that drive new gene evolution, gene fusion has emerged as a significant yet relatively understudied contributor. It can arise through chromosomal rearrangements or RNA processing mechanisms, merging segments from different genes to produce novel fusion transcripts. In plants, these fusion events have been associated with key biological functions, including the regulation of specialized metabolism, stress responses, and developmental changes. While fusion genes have been extensively studied in humans, mainly due to their oncogenic potential, their prevalence and functional relevance in plants remain relatively underexplored. This review offers a detailed overview of the molecular mechanisms underlying gene fusion formation, highlighting their participation in gene evolution, functional diversification, and plant adaptation. In addition, we discuss current methodologies for detecting and validating fusion events, including high-throughput sequencing technologies and emerging single-cell sequencing platforms, and outline promising directions for future research aimed at elucidating their biological significance. Collectively, these insights emphasize the expanding importance of gene fusions in plant biology and underscore the need for further investigation into their regulatory and evolutionary roles.Item A CRISPR-Cas9 library to target putative redundant gene sets facilitates their functional exploration in grain development in rice(Springer Nature Publishing AG, 2025) Yadav, Banita; Sardar, Shaswati; Yadav, Anil; Kumari, Annapurna; Gautam, Mohini; Mandlik, Rushil; Arora, Simran; Kumar, Shailesh; Jewaria, Pawan Kumar; Sonah, Humira; Deshmukh, Rupesh; Chinnusamy, Viswanathan; Ram, HasthiAdvent of CRISPR-Cas9 library approach has revolutionized the field of high throughput targeted mutagenesis in plants. By identifying an sgRNA spacer that can target multiple paralogous genes in a genome, higher-order knockout plants can be developed. Using this concept, we developed ten CRISPR-Cas9 pool libraries and generated higher-order knockout plants in rice. Towards this, firstly we identified genome-wide sets of genes which are co-expressed and have high sequence similarity and can be targeted by a single sgRNA. Based on the expression pattern, these genes were divided into ten groups, and subsequently ten CRISPR-Cas9 plasmid libraries were developed. One such library designed against seed-expressed genes was transformed into rice and higher-order knockout plants were developed. Genotyping revealed that around 90% T0 plants had editing, and among the edited plants majority of them were higher-order knockouts. Phenotypic analysis in the next generation discovered functions of several seed specific genes in grain length, width, number and 100-grain weight. By analyzing single and double mutants for two Agenet domain-containing proteins, we have discovered an epistatic interaction between them for grain development. Further application of our approach will help to uncover hidden functions of the targeted genes and accelerate functional genomics research in rice. The CRISPR-Cas9 library is a useful approach to generate higher-order knockout mutants and identify functions of the targeted genes in rice.Item Molecular and expression analyses indicate the role of fusion transcripts in mediating abiotic stress responses in chickpea(Frontiers Media S.A., 2025) Hamid, Fiza; Zahra, Shafaque; Kumar, ShaileshUnderstanding the transcriptome diversity is essential for deciphering the transcriptional level regulation. High-throughput sequencing technologies have facilitated the detection of fusion transcripts (FTs), which are chimeric mRNA molecules derived from gene fusions due to chromosomal rearrangements or via the splicing machinery at the RNA level. In this study, we investigated the transcriptome complexity in Cicer arietinum resulting from fusion events using high-throughput RNA-Seq datasets from five tissues, i.e., stem, leaves, buds, flowers, and pods, and two abiotic stress conditions, i.e., drought and salinity. Of the 328 unique FTs identified, 69% exhibited the presence of canonical splice sites at their junction, indicating their generation via trans-splicing. Functional annotation and enrichment analyses of fusion partners suggested that these transcripts may expand functional diversity. A total of 10 FTs were validated via RT-PCR followed by Sanger sequencing, which are the first FTs described in the important legume chickpea. Expression analysis of fusion transcripts across various tissues and under abiotic stress conditions revealed evidence of context-dependent regulation. Furthermore, 120 fusion gene pairs were found to be conserved across 17 chickpea genotypes, highlighting their potential biological significance and stability within the species. Overall, these findings suggest that fusion transcripts may contribute to regulatory mechanisms underlying abiotic stress responses in chickpea.Item The RNA-binding protein Quaking is essential for cardiac homeostasis and function by regulating Morf4l2 splicing(Elsevier B.V., 2026) Kumari, Sunaina; Shashi; Singh, Sandhya; Swain, Abinash; Prakash, Shakti; Chitkara, Pragya; Sharma, Rakesh Kumar; Agarwal, Pratyush; Kundu, Samprikta; Gaur, Aakash; Kumari, Renu; Sinha, Abhipsa; Chatterjee, Shambhabi; Prasun, Pankaj; Hummel, Oliver; Pant, Bhaskar; Srivastava, Kinshuk Raj; Hübner, Norbert; Datta, Dipak; Mitra, Kalyan; Mishra, Durga Prasad; Guha, Rajdeep; Thum, Thomas; Kumar, Shailesh; Gupta, Shashi KumarBackground: Lower levels of Qki were reported in human and mouse-failing hearts, implicating its involvement in cardiac diseases. However, the molecular and functional effects of its downregulation in adult myocardium remain largely unknown. Objective: We aim to uncover the effects of Qki knockdown in adult hearts. Methods & results: Here we show that AAV9-mediated knockdown of Qki by shRNAs in the hearts of adult BALB/c mice led to cardiac malfunction, atrophy, apoptosis, heart failure, and death within two weeks. Global transcriptomic analysis of Qki knockdown hearts revealed significant dysregulation of 996 alternative splicing events upon Qki knockdown. Mechanistically, we discovered that loss of Qki promotes the exclusion of the third exon of Morf4l2, leading to higher expression of exon three excluded variant (Morf4l2Δex3). Like rodents, the RNA-seq dataset from 108 human hearts revealed a lower splice junction count of MORF4L2 exon three in hearts with low levels of QKI compared to subjects with higher QKI levels. Specific knockdown of Morf4l2Δex3 rescues Qki knockdown-induced cardiac cachexia and improves cardiac function. Moreover, Morf4l2Δex3 was increased in the colon cancer-induced cardiac cachexia mouse model, and its inhibition prevented cardiac cachexia and improved cardiac function. Mechanistically, exon three of Morf4l2 lies in the 5'UTR, and its exclusion leads to higher expression of MORF4L2 upon Qki knockdown due to the lack of a G2-quadruplex. Importantly, MORF4L2 protein sequence and localization were not affected by alternative splicing as exon three lies in the 5'UTR. We found that MORF4L2 is a chromatin-bound protein and regulates H3K27ac. Conclusion: Qki knockdown in the adult heart leads to cardiac cachexia due to the alteration of Morf4l2 splicing. Inhibition of Morf4l2Δex3 inhibits cancer-induced cardiac cachexia, demonstrating it as a potential therapeutic target.
