Browsing by Author "Mukherjee, Kanka"
Now showing 1 - 4 of 4
- Results Per Page
- Sort Options
Item AquaaG: A comprehensive pipeline for quality assessment and annotation of genomes(Elsevier B.V., 2026) Shukla, Jagriti; Mukherjee, Kanka; Sahu, Namrata; Kumar, ShaileshThe rapid expansion of publicly available genome assemblies has made genome annotation an increasingly challenging task, particularly for large-scale analyses across prokaryotic and eukaryotic organisms. While several tools exist for assembly evaluation and annotation, their use often involves fragmented workflows that require extensive manual coordination. To overcome this limitation, we introduce AquaaG, an automated and reproducible genome annotation pipeline. AquaaG integrates genome assembly retrieval from NCBI, assembly quality assessment using QUAST, organism-specific annotation using Prokka for prokaryotes and BRAKER3 for eukaryotes, gene-space completeness evaluation using BUSCO, and functional annotation using EggNOG-mapper. The pipeline is configured through simple YAML files and supports species-level, kingdom-level, and custom assembly-based analyses with optional submitter-based filtering. AquaaG therefore provides a practical and reproducible framework for high-throughput genome annotation and assessment.Item AraNSdb: a dedicated database of stress-responsive non-coding RNAs in Arabidopsis thaliana(Springer Nature Publishing AG, 2026) Vivek, A.T.; Bhatia, Manika; Sahu, Namrata; Kalakoti, Garima; Kaushik, Love; Mukherjee, Kanka; Kumar, ShaileshPlants, as sessile organisms, are constantly exposed to biotic and abiotic stresses, making their ability to respond crucial for survival. Non-coding RNAs (ncRNAs) have emerged as key regulators in these stress responses, with several studies identifying numerous stress-responsive ncRNAs (SRNs). However, a comprehensive collection of SRNs derived from sequencing data in Arabidopsis thaliana has been lacking. To address this, we utilized high-throughput experimental data and mined published studies to construct AraNSdb (Arabidopsis ncRNA Stress Database), a systematic resource for storing and querying SRNs. AraNSdb documents over 1,000 expression profiles from diverse stress datasets, encompassing 6,616 SRNs, including microRNAs (miRNAs), small interfering RNAs (siRNAs), long non-coding RNAs (lncRNAs), and circular RNAs (circRNAs). The database features an intuitive web interface for exploring SRNs associated with specific stress types and provides detailed ncRNA annotations to support functional and regulatory studies. AraNSdb offers a valuable platform for advancing our understanding of ncRNA-mediated stress responses and is freely accessible at http://www.nipgr.ac.in/AraNSdb.Item PFGPred: A stack ensemble classifier for the identification of fusion genes in plants(Oxford University Press, 2026) Hamid, Fiza; Mukherjee, Kanka; Chaudhary, Sakshi; Kaushik, Love; Kumar, ShaileshFusion genes play crucial roles in plant biological processes but remain far less explored than their human counterparts, largely due to limited validated datasets and the absence of plant-specific prediction tools. Existing approaches often produce high false-positive rates, restricting reliable discovery. To address this gap, we developed Plant Fusion Gene Predictor (PFGPred), an ensemble machine learning framework that integrates Random Forest, XGBoost, and long short-term memory (LSTM) models into a meta-classifier for accurate identification of true and false fusion genes from RNA-Seq data. PFGPred was trained on a high-confidence dataset of fusion genes validated by both RNA-Seq and whole-genome sequencing from Arabidopsis thaliana, Oryza sativa, Triticum aestivum, and Zea mays, to predict and rank candidate fusion genes for future functional validation. It outperformed individual baseline models, achieving accuracies of 0.97 on training data and 0.77 on independent test data. When evaluated on human datasets, it achieved 0.71 accuracy with lower sensitivity, reflecting biological differences between plant and human fusion events. Comparative analyses confirmed that PFGPred reliably identifies validated fusions, demonstrating its utility as a cost-effective, plant-specific prediction tool for high-throughput fusion gene screening and functional genomics research. It is freely available as a web server at http://www.nipgr.ac.in/PFGPred.Item The intersection of AI and genomics in health and disease: Advancements and applications(Elsevier B.V., 2026) Kaushik, Love; Vivek, A T; Arora, Simran; Hamid, Fiza; Mukherjee, Kanka; Bisht, Niyati; Chaudhary, Sakshi; Shukla, Jagriti; Nawani, Sakshi; Kumar, ShaileshAI and genomics are revolutionizing precision medicine by using machine learning (ML) to analyze large-scale next-generation sequencing (NGS) data, identifying genetic mutations and biomarkers for personalized therapies. In practice, this accelerates drug discovery and enhances variant detection, while in cancer genomics, AI enables early detection via liquid biopsies and refines treatment by integrating multi-omics data to improve therapeutic precision. However, challenges such as data biases in underrepresented populations, limited model interpretability, and ethical concerns regarding privacy and algorithmic inequity hinder clinical adoption and demand robust governance. Efforts to diversify datasets also face standardization hurdles, although explainable AI and federated learning provide promising solutions for improving transparency and privacy. In this chapter, we discuss the role of AI in advancing genomics from diagnostics to novel therapies and emphasize the need for equitable frameworks to ensure responsible implementation, thereby paving the way for breakthroughs in personalized medicine.
