Publications of NIPGR Scientists

Permanent URI for this communityhttps://ndkr-library.nipgr.ac.in/handle/123456789/1

Browse

Search Results

Now showing 1 - 9 of 9
  • Item
    The intersection of AI and genomics in health and disease: Advancements and applications
    (Elsevier B.V., 2026) Kaushik, Love; Vivek, A T; Arora, Simran; Hamid, Fiza; Mukherjee, Kanka; Bisht, Niyati; Chaudhary, Sakshi; Shukla, Jagriti; Nawani, Sakshi; Kumar, Shailesh
    AI and genomics are revolutionizing precision medicine by using machine learning (ML) to analyze large-scale next-generation sequencing (NGS) data, identifying genetic mutations and biomarkers for personalized therapies. In practice, this accelerates drug discovery and enhances variant detection, while in cancer genomics, AI enables early detection via liquid biopsies and refines treatment by integrating multi-omics data to improve therapeutic precision. However, challenges such as data biases in underrepresented populations, limited model interpretability, and ethical concerns regarding privacy and algorithmic inequity hinder clinical adoption and demand robust governance. Efforts to diversify datasets also face standardization hurdles, although explainable AI and federated learning provide promising solutions for improving transparency and privacy. In this chapter, we discuss the role of AI in advancing genomics from diagnostics to novel therapies and emphasize the need for equitable frameworks to ensure responsible implementation, thereby paving the way for breakthroughs in personalized medicine.
  • Thumbnail Image
    Item
    Breaking and making genes: the genesis of novel traits in plants
    (John Wiley & Sons, 2026) Hamid, Fiza; Arora, Simran; Kumar, Shailesh
    Understanding the mechanisms by which plants adapt, evolve, and acquire new traits is crucial for enhancing agricultural resilience and productivity in the face of global challenges. Among the various mechanisms that drive new gene evolution, gene fusion has emerged as a significant yet relatively understudied contributor. It can arise through chromosomal rearrangements or RNA processing mechanisms, merging segments from different genes to produce novel fusion transcripts. In plants, these fusion events have been associated with key biological functions, including the regulation of specialized metabolism, stress responses, and developmental changes. While fusion genes have been extensively studied in humans, mainly due to their oncogenic potential, their prevalence and functional relevance in plants remain relatively underexplored. This review offers a detailed overview of the molecular mechanisms underlying gene fusion formation, highlighting their participation in gene evolution, functional diversification, and plant adaptation. In addition, we discuss current methodologies for detecting and validating fusion events, including high-throughput sequencing technologies and emerging single-cell sequencing platforms, and outline promising directions for future research aimed at elucidating their biological significance. Collectively, these insights emphasize the expanding importance of gene fusions in plant biology and underscore the need for further investigation into their regulatory and evolutionary roles.
  • Thumbnail Image
    Item
    A CRISPR-Cas9 library to target putative redundant gene sets facilitates their functional exploration in grain development in rice
    (Springer Nature Publishing AG, 2025) Yadav, Banita; Sardar, Shaswati; Yadav, Anil; Kumari, Annapurna; Gautam, Mohini; Mandlik, Rushil; Arora, Simran; Kumar, Shailesh; Jewaria, Pawan Kumar; Sonah, Humira; Deshmukh, Rupesh; Chinnusamy, Viswanathan; Ram, Hasthi
    Advent of CRISPR-Cas9 library approach has revolutionized the field of high throughput targeted mutagenesis in plants. By identifying an sgRNA spacer that can target multiple paralogous genes in a genome, higher-order knockout plants can be developed. Using this concept, we developed ten CRISPR-Cas9 pool libraries and generated higher-order knockout plants in rice. Towards this, firstly we identified genome-wide sets of genes which are co-expressed and have high sequence similarity and can be targeted by a single sgRNA. Based on the expression pattern, these genes were divided into ten groups, and subsequently ten CRISPR-Cas9 plasmid libraries were developed. One such library designed against seed-expressed genes was transformed into rice and higher-order knockout plants were developed. Genotyping revealed that around 90% T0 plants had editing, and among the edited plants majority of them were higher-order knockouts. Phenotypic analysis in the next generation discovered functions of several seed specific genes in grain length, width, number and 100-grain weight. By analyzing single and double mutants for two Agenet domain-containing proteins, we have discovered an epistatic interaction between them for grain development. Further application of our approach will help to uncover hidden functions of the targeted genes and accelerate functional genomics research in rice. The CRISPR-Cas9 library is a useful approach to generate higher-order knockout mutants and identify functions of the targeted genes in rice.
  • Item
    Identification of tRNA-derived fragments in legumes
    (Springer Nature Publishing AG, 2026) Arora, Simran; Aftab, Sahrish; Shree, Tanu; Kumar, Shailesh
    The tRNA-derived noncoding RNAs (tncRNAs) belong to the novel class of noncoding RNAs, acting as important components of genome regulatory circuits. In planta, the mechanism of generation and function of tncRNAs is not fully elucidated. Production of important leguminous plants like chickpea, Medicago and soybean is majorly hampered due to different biotic and abiotic stresses. Identification and characterization of tncRNAs in legumes may open a new paradigm for molecular biologists to make novel tools for the improved varieties of legumes for sustainable agriculture. The first step in the study of tncRNAs is to identify and annotate them in small RNA sequencing datasets. Here, we have demonstrated the tncRNA Toolkit for the identification and annotation of tncRNAs in a small RNA sequencing dataset of the important legume crop chickpea.
  • Thumbnail Image
    Item
    Fusion transcripts in plants: hidden layer of transcriptome complexity
    (Elsevier B.V., 2025) Arora, Simran; Hamid, Fiza; Kumar, Shailesh
    In the realm of genetic information, fusion transcripts contribute to the intricate complexity of the transcriptome across various organisms. Recently, Cong et al. investigated these RNAs in rice, maize, soybean, and arabidopsis (Arabidopsis thaliana), revealing conserved characteristics. These findings enhance our understanding of the functional roles and evolutionary significance of these fusion transcripts.
  • Thumbnail Image
    Item
    The landscape of fusion transcripts in plants: a new insight into genome complexity
    (BioMed Central Ltd, 2024) Chitkara, Pragya; Singh, Ajeet; Gangwar, Rashmi; Bhardwaj, Rohan; Zahra, Shafaque; Arora, Simran; Hamid, Fiza; Arya, Ajay; Sahu, Namrata; Chakraborty, Srija; Ramesh, Madhulika; Kumar, Shailesh
    Background Fusion transcripts (FTs), generated by the fusion of genes at the DNA level or RNA-level splicing events significantly contribute to transcriptome diversity. FTs are usually considered unique features of neoplasia and serve as biomarkers and therapeutic targets for multiple cancers. The latest findings show the presence of FTs in normal human physiology. Several discrete reports mentioned the presence of fusion transcripts in planta, has important roles in stress responses, morphological alterations, or traits (e.g. seed size, etc.). Results In this study, we identified 169,197 fusion transcripts in 2795 transcriptome datasets of Arabidopsis thaliana, Cicer arietinum, and Oryza sativa by using a combination of tools, and confirmed the translational activity of 150 fusion transcripts through proteomic datasets. Analysis of the FT junction sequences and their association with epigenetic factors, as revealed by ChIP-Seq datasets, demonstrated an organised process of fusion formation at the DNA level. We investigated the possible impact of three-dimensional chromatin conformation on intra-chromosomal fusion events by leveraging the Hi-C datasets with the incidence of fusion transcripts. We further utilised the longread RNA-Seq datasets to validate the most reoccurring fusion transcripts in each plant species followed by further authentication through RT-PCR and Sanger sequencing. Conclusions Our findings suggest that a significant portion of fusion events may be attributed to alternative splicing during transcription, accounting for numerous fusion events without a proportional increase in the number of RNA pairs. Even non-nuclear DNA transcripts from mitochondria and chloroplasts can participate in intra- and inter-chromosomal fusion formation. Genes in close spatial proximity are more prone to undergoing fusion formation, especially in intra-chromosomal FTs. Most of the fusion transcripts may not undergo translation and serve as long non-coding RNAs. The low validation rate of FTs in plants indicated that the fusion transcripts are expressed at very low levels, like in the case of humans. FTs often originate from parental genes involved in essential biological processes, suggesting their relevance across diverse tissues and stress conditions. This study presents a comprehensive repository of fusion transcripts, offering valuable insights into their roles in vital physiological processes and stress responses.
  • Thumbnail Image
    Item
    PFusionDB: a comprehensive database of plant-specific fusion transcripts
    (Springer Nature Publishing AG, 2024) Arya, Ajay; Arora, Simran; Hamid, Fiza; Kumar, Shailesh
    Fusion transcripts (FTs) are well known cancer biomarkers, relatively understudied in plants. Here, we developed PFusionDB (www.nipgr.ac.in/PFusionDB), a novel plant-specific fusion-transcript database. It is a comprehensive repository of 80,170, 39,108, 83,330, and 11,500 unique fusions detected in 1280, 637, 697, and 181 RNA-Seq samples of Arabidopsis thaliana, Oryza sativa japonica, Oryza sativa indica, and Cicer arietinum respectively. Here, a total of 76,599 (Arabidopsis thaliana), 35,480 (Oryza sativa japonica), 72,099 (Oryza sativa indica), and 9524 (Cicer arietinum) fusion transcripts are non-recurrent i.e., only found in one sample. Identification of FTs was performed by using a total of five tools viz. EricScript-Plants, STAR-Fusion, TrinityFusion, SQUID, and MapSplice. At PFusionDB, available fundamental details of fusion events includes the information of parental genes, junction sequence, expression levels of fusion transcripts, breakpoint coordinates, strand information, tissue type, treatment information, fusion type, PFusionDB ID, and Sequence Read Archive (SRA) ID. Further, two search modules: ‘Simple Search’ and ‘Advanced Search’, along with a ‘Browse’ option to data download, are present for the ease of users. Three distinct modules viz. ‘BLASTN’, ‘SW Align’, and ‘Mapping’ are also available for efficient query sequence mapping and alignment to FTs. PFusionDB serves as a crucial resource for delving into the intricate world of fusion transcript in plants, providing researchers with a foundation for further exploration and analysis. Database URL: www.nipgr.ac.in/PFusionDB.
  • Item
    A protocol for the detection of fusion transcripts using RNA-sequencing data
    (Springer Nature Publishing AG, 2024) Hamid, Fiza; Arora, Simran; Chitkara, Pragya; Kumar, Shailesh
    Fusion transcripts are formed when two genes or their mRNAs fuse to produce a novel gene or chimeric transcript. Fusion genes are well-known cancer biomarkers used for cancer diagnosis and as therapeutic targets. Gene fusions are also found in normal physiology and lead to the evolution of novel genes that contribute to better survival and adaptation for an organism. Various in vitro approaches, such as FISH, PCR, RT-PCR, and chromosome banding techniques, have been used to detect gene fusion. However, all these approaches have low resolution and throughput. Due to the development of high-throughput next-generation sequencing technologies, the detection of fusion transcript becomes feasible using whole genome sequencing, RNA-Seq data, and bioinformatics tools. This chapter will overview the general computational protocol for fusion transcript detection from RNA-sequencing datasets.
  • Item
    In silico identification of tRNA fragments, novel candidates for cancer biomarkers, and therapeutic targets
    (Springer Nature Publishing AG, 2024) Singh, Ankita; Zahra, Shafaque; Arora, Simran; Hamid, Fiza; Kumar, Shailesh
    The identification of a wide variety of RNA molecules using high-throughput sequencing techniques in the transcriptome pool of living organisms has revealed hidden regulatory insights in the cell. The class of non-coding RNA fragments produced from transfer RNA, or tRFs, is one such example. They are heterogeneously sized molecules with lengths ranging between 15 and 50 nt. They have a history of being dysregulated in human malignancies and other illnesses. The detection of these molecules has been made easier by a variety of bioinformatics techniques. The various types of tRFs and how they relate to cancer are covered in this chapter. It also provides a summary of the biological significance of tRFs reported in human cancer. Additionally, it emphasizes the utilities of databases and computational tools that have been created by different research teams for the investigation of tRFs. This will further aid the exploration and analysis of tRFs in cancer research and will support future advancement and a better comprehension of these molecules.