Experience

Professional

Current
Data Scientist, CSG CTO Lab - Dell Technologies Bengaluru ยท Jul 2025 โ€“ Present

What I work on right now:

  • Efficient on-device AI - getting large models to run well on memory-constrained consumer hardware, rather than offloading everything to the cloud.
  • Hardware-aware optimization - learning-based scheduling of AI workloads across heterogeneous silicon, trading off latency, energy, and quality.
  • Reliability under compression - keeping models calibrated, factually grounded, and secure as they get aggressively quantized.

Published work from this role:

  • Engineered StreamServe (arXiv:2604.09562), a prefill-decode disaggregated LLM serving system co-optimizing multi-signal routing (FlowGuard) and dynamic speculative execution (SpecuStream): 11-18ร— latency reduction and up to 4.4ร— average throughput over TP-vLLM baselines on 4ร— A800 GPUs.
  • Developed RAMP (arXiv:2603.17891), the first zero-shot transferable mixed-precision quantization policy for LLMs: bit allocation as a constrained MDP solved via Soft Actor-Critic, with Scale Folding to absorb activation outliers. SOTA Pareto frontier at 3.65 effective bits (3.68 GB) at 5.54 perplexity, outperforming AWQ/GPTQ, with zero-shot transfer to Mistral-7B and Llama-2-13B.
  • Designed DiffuTruth (FEVER @ EACL 2026), an unsupervised hallucination detector grounded in non-equilibrium thermodynamics: a Generative Stress Test via discrete text diffusion computes NLI-based semantic energy. SOTA 0.725 AUROC on FEVER and a >4% zero-shot robustness gain on multi-hop HOVER.
LLM InferenceDisaggregated ServingQuantizationRLHallucination Detection
Previous
Data Science Intern - Dell Technologies Bengaluru ยท Jul 2024 โ€“ Jun 2025
  • Developed CogniSQL-R1-Zero (arXiv:2507.06013), an execution-aligned Text-to-SQL reasoning model trained via pure GRPO on a lightweight sparse-reward signal with no cold-start SFT; converged in 6.2 hours on 4ร— A100 GPUs via DeepSpeed ZeRO-2.
  • Set SOTA for mid-sized models: 59.97% execution accuracy on BIRD with a 7B backbone, outperforming Mistral 123B, DeepSeek-Coder 236B, and GPT-4; pushed to 69.68% via test-time scaling with best-of-6 execution search.
  • Architected a parallel multi-agent reasoning pipeline (query decomposition, runtime self-healing, majority-voting ensembles): +30% execution accuracy on proprietary schemas, shipped as a deployed enterprise Text-to-SQL copilot.
GRPOText-to-SQLDeepSpeedAgentic AI
Computer Vision Research Engineer - Stealth Startup (Remote, US) Mar 2024 โ€“ Jun 2024
Real-time theft detection using SlowFast networks and 3D CNNs, optimized for low-latency edge deployment.
Computer VisionSlowFastEdge AI
AI Research Intern - Renix Informatics Aug 2023 โ€“ Oct 2023
Researched and implemented two features for the DocX Document AI product. Used SonarQube for code quality analysis.
AI/ML Developer + Mentoring Intern - WictroniX Jun 2023 โ€“ Aug 2023
AI/ML work on drone footage (120m altitude). Led mentoring group guiding 30+ interns.
AI/ML Apprentice - IBM Z Systems Jan 2023 โ€“ Jun 2023
Developed two end-to-end ML projects (Omnizenon & KnowCrimez). Applied ML for data analysis and deployed web UIs.
Machine Learning Intern - Suvidha Foundation Mar 2023 โ€“ Apr 2023
Abstractive text summarization using HuggingFace Transformers. Benchmarked accuracy across LLMs.
Backend Developer Trainee - Safcurl Jul 2022 โ€“ Jan 2023
Sprint tickets, model building, API implementation, and database maintenance.

Patents

Federated and Self-Learning Techniques for Root Cause Detection in Edge-Cloud Environments Filed May 2026
Co-inventor. U.S. Patent Application No. 19/670,270 (pending).
6 Edge-AI inventions approved for USPTO filing 2025 โ€“ 2026
Co-inventor. Approved by the Dell Technologies Internal Patent Committee; formal drafts in progress with Dell Legal.

Research

Research Collaborator - Manipal University Jaipur Remote ยท Jul 2025 โ€“ Present
Advisor: Dr. Varun Tiwari.
  • D-CLS (under review, Complex & Intelligent Systems): a training-free logit-suppression decoding method that mitigates object hallucination in VLMs, cutting CHAIR hallucination metrics by up to 19.1% at zero compute cost while preserving fluency.
  • Amortization as Robustness (under review, Neural Computing and Applications): a distributionally robust RL framework (PPO) for algorithmic recourse, achieving 2.6ร— lower L2 cost and 28ร— lower deployment latency across heterogeneous Rashomon ensembles.
  • PRISM (under review, Discover Computing): a mechanistic diagnostic framework using component-wise mixed precision, showing that low-bit quantization disproportionately collapses MLP parametric memory while preserving external RAG context utilization.
Undergraduate Researcher - Manipal University Jaipur Feb 2024 โ€“ Jun 2025
  • Hybrid CNN+GRU+LSTM for lymphoma detection from histopathology images - Dr. Vivek Bhardwaj (accepted ICCCNT @ IIT Indore 2025, oral).
  • Real-time QUIC traffic classifier with LightGBM and SHAP+LIME explainability - Mr. Rajesh Kumar (accepted ICCCNT @ IIT Indore 2025).
Winter Research Intern - VECC, Dept. of Atomic Energy, Kolkata Dec 2023 โ€“ Feb 2024
RL (TD3)-based autonomous navigation for robots in nuclear radiation environments - Ushnish Sarkar (Scientific Officer F).
Summer Research Intern - IIT BHU, Varanasi Jun 2023 โ€“ Jul 2023
Hyperspectral image classification on the QUH dataset (10โถ+ samples) - Prof. Rajeev Srivastava. Hybrid deep learning model achieving 91.90% accuracy.

Education

Manipal University Jaipur, India Oct 2021 โ€“ Jun 2025
B.Tech. (Hons) CSE - Specialization: AI & ML ยท GPA 8.56/10
Dean's List 7ร— ยท 4ร— Student Excellence Award ยท MUJ Wizard Programmer Gold Medal ยท Tea with President Award

Teaching & Volunteering

Subject Matter Expert - IBM, Global Oct 2023 โ€“ Jun 2024
Led workshops at IBM Z Datathon guiding 3300+ students in LinuxONE for AI development. Mentored winning teams on mainframe-based ML integration.
Teaching Assistant - Manipal University Jaipur Jan 2024 โ€“ May 2024
Tutorials and lab sessions for AI3241 Reinforcement Learning (Dr. Animesh Kumar) and AI3231 Computer Vision & Pattern Lab (Prof. Harish Sharma).
Open-Source Contributor - Deep-ML 2025
Author of educational challenges on Deep-ML spanning quantization, RAG, RLHF, and inference optimization.
Hackathon Mentor (4ร—) 2023 โ€“ 2025
Mentored teams at global and regional hackathons.
President & Co-Founder - AIML Community MUJ Sep 2023 โ€“ Ongoing
300+ members in first year. Organized events and hackathons on ML, CV, and NLP.
IBM Z Student Ambassador - IBM Z Systems Aug 2023 โ€“ Ongoing
Speaker at IBM Z Day. Promoting mainframe technology globally.
Student Ambassador Leader - Streamlit Aug 2023 โ€“ Ongoing
One of 10 leaders among 216 ambassadors globally. Works with the Streamlit team directly.
Deputy Head of Design - Phi Phenomenon MUJ Jun 2022 โ€“ Jul 2023

Skills

Languages & Frameworks:

PythonCUDATritonPyTorch TensorFlowTRLDeepSpeedLangChain OpenCVScikit-learnHuggingFace StreamlitJavaSQLC++

LLM Serving & Quantization:

vLLMTensorRT-LLMGPTQAWQ SmoothQuantNF4MXFP4 KV-cache compressionSpeculative decoding

RL & Training:

SACPPOGRPODPO Knowledge distillationRLHF

Interpretability & Evaluation:

TransformerLensnnsightActivation patching Linear probesCalibration (ECE)Selective prediction PySR

Core Competencies:

LLM InferenceQuantizationDistributed Training Reinforcement LearningRAGReasoning LLMs Mechanistic InterpretabilityAI Safety Computer VisionNLPAgentic AI Mainframes / IBM ZRobotics