I build reinforcement learning systems, autonomous robotics, and applied machine learning — from constrained decision-making frameworks to physics-informed digital twins and multi-agent AI. Most of my work is public and reproducible.
Open-source physics-informed digital twin for turbojet engines: health monitoring, prognostics, explainable AI, and interactive 3D visualization.
Repository ↗End-to-end AO pipeline: SH-WFS centroiding with 127× C speedup → CNN/MMSE reconstruction (Strehl 0.9987) → LQG control → SLODAR profiling. BAH 2026 Challenge #9.
Repository ↗IMM-EKF multi-target tracking, APN guidance, Hungarian-assignment coordination. Validated at 88.7% success rate with 13ms latency on Raspberry Pi 5 + Arduino.
Project site ↗Constrained RL framework for delayed, state-conditioned penalties. Delay-corrected Bellman operator with proven contraction property. Submitted to NeurIPS 2026.
Repository ↗
I'm from Patna, Bihar, working independently on reinforcement learning, robotics, and applied machine learning. My work spans constrained RL frameworks, sensor fusion and tracking, digital twins, multi-agent AI, and embedded systems. I build things end-to-end — from physics models and training code to dashboards and hardware.
I'm also founder and CEO of D-MechatronicX, a robotics and defense-AI team (company registration currently in process) — the team behind the Technoxian and Robotex results below.
I write my own code, run my own benchmarks, and publish what I can verify. Where a result is simulated rather than field-tested, I say so directly — that distinction matters more to me than how impressive a number sounds.
I'm looking for research collaborators and mentorship, particularly in reinforcement learning, estimation theory, and multi-agent systems, and I'm open to feedback on any of the work below.
Open-source physics-informed digital twin platform for turbojet engines: engine health monitoring, prognostics, explainable AI, and interactive 3D visualization. Combines thermodynamic modeling with machine learning for real-time diagnostics.
End-to-end adaptive optics pipeline: Shack-Hartmann WFS centroiding with 127× C speedup → CNN/MMSE wavefront reconstruction (Strehl 0.9987) → LQG closed-loop control → SLODAR turbulence profiling. BAH 2026 Challenge #9.
Autonomous drone interception system with IMM-EKF multi-target tracking, augmented proportional navigation (APN) guidance, and Hungarian-assignment multi-agent coordination. Validated at 88.7% success rate with 13ms latency on Raspberry Pi 5 + Arduino hardware.
A reinforcement learning framework, implemented in pure NumPy, for constrained decision-making where an action's cost is state-conditioned and delayed. Estimates a latent delay term and applies a state-conditioned penalty rather than fixed reward shaping. Currently v6.0. Submitted to NeurIPS 2026.
Local multi-agent AI system with 6 specialized agents, 12 tools, ChromaDB vector memory, and Ollama LLM backend. Fully offline capable with tool use, memory, and inter-agent coordination.
Co-founded with Jay Aditya and Hardit Singh; I lead the technical side as CTO. Syntheta multiplies real-world capture into independently labeled training variants for physical AI (factories, robots, drones): a single frame is restored for low light, then multiplied via domain-randomized mutation, neural view-synthesis, and a clean passthrough baseline. On its last validated run, the pipeline produced roughly 9.3 output frames per input frame, with 128 passing unit tests.
A 3D-reconstruction extension (structure-from-motion into Gaussian splats) is in progress and explicitly incomplete. Applied to Y Combinator, decision pending; pre-revenue and not yet incorporated.
An independent study implementing a classical multi-target tracking and guidance pipeline in simulation: IMM-EKF for maneuvering-target tracking, augmented proportional navigation (APN) guidance, and Hungarian-algorithm multi-agent assignment. The goal was to understand how these three classical techniques compose.
All performance figures come from Monte Carlo simulation on synthetic trajectories, not physical hardware trials or field tests.
A simulation applying CCPL to hospital bed and staff allocation under emergency surge conditions, modeling patient acuity with a Kalman-filter-style tracker. A research simulation, not a system used by any hospital or in any clinical setting.
A progressive web app for road emergencies, built for the IIT Madras Road Safety Hackathon 2026. Detects crash impacts via the device accelerometer, broadcasts location to emergency contacts and services (108/100/101/112 shortcuts), includes an on-device medical ID, golden-hour countdown timer, community hazard reporting, and an AI-assisted first-aid guide. Interface is available in Hindi and five regional Indian languages, and core safety functions work offline.
An intelligent Raspberry Pi-powered robotic assistant featuring voice interaction, memory, expressive OLED animations, autonomous hardware control, and modular AI architecture inspired by TARS from Interstellar.
Production-grade bird vocalization classifier that identifies species and decodes communication meaning (alarm, mating, territorial...) using EfficientNet-B0 on mel-spectrograms, served via FastAPI and containerized with Docker.
Deep learning system that classifies infant cry reasons (hunger, pain, tiredness...) and age group from raw audio using EfficientNet-B0 on mel-spectrograms. Designed as a assistive tool for caregivers and healthcare workers.
Building a digital organism that learns, grows, and develops intelligence through lifelong experience. An early-stage research project exploring embodied cognition and continuous learning in silico.
Turns whatever you already have — a notebook, a WhatsApp group, a voice update — into a live twin of your factory. Then it tells you, in plain COO language, which decisions today move ₹ the most.
Multi-agent AI civilization simulator that runs in your terminal. Guide AI agents from the Stone Age to the Iron Age, make council decisions, survive disasters, discover technologies, and build a legacy. No browser required.
A local tool that evaluates hackathon submissions — slide decks, source code, CAD files, and GitHub repositories — using an open-weight LLM (Ollama/Llama 3) run entirely offline, plus static analysis (AST-based code metrics, mesh inspection for CAD). Built to remove dependence on paid API access when judging submissions.
A personal robotics control environment built on Electron: a subsystem scheduler, sensor and actuator management dashboards, and a basic 2D navigation map, intended as a testbed for coordinating multiple robotics subsystems from one interface rather than a production robot OS.
A wearable personal-safety device combining GPS location tracking with GSM-based SOS alerting to pre-set emergency contacts. Recognized in the India Book of Records; published as first-author research in IJNRD (see Awards).
An Arduino-based mechanism using an ultrasonic sensor and servo motor to lift a face mask automatically on hand approach, built during COVID-19 restrictions. Total component cost was approximately ₹280.
A “tamagotchi for programmers” — conversational AI that remembers context and interacts with you over time. Built as a personal experiment in persistent AI personality and memory.
Build causal timelines, fork realities, and simulate the roads not taken. A creative simulation engine for exploring counterfactual history and branching causality.
Frontend bug fixes and improvements for Kartzia Eeden e-commerce app: fixed stale cart total, session restore, form submission no-op, auth guards, gradient styling, and accessibility issues.
Scheduling infrastructure for absolutely everyone. A community-driven calendar and scheduling platform.
A reinforcement learning paper addressing three problems in how constrained RL methods compute safety penalties: unknown causal delay between action and consequence, conflation of an agent's own effects with consequences already in motion, and non-stationary Bellman targets when the penalty weight is updated. Proposes a delay-corrected Bellman operator with a proven contraction property, a state-conditioned penalty shown to dominate scalar penalties, and a causal-contribution estimator (an Interventional Consequence Net) trained on ground-truth labels from a structural causal model.
Submitted to NeurIPS 2026 and currently under review — this is not yet an acceptance, and OpenReview forum pages for this track are not public until decisions are released. Submission number and abstract are drawn directly from the author's own OpenReview submission record.
Open to research collaboration, mentorship conversations, and feedback on any of the work above. I reply to email.