I build software that works in the physical world: a counter-drone defense engine, an open-source RL framework for safety under delayed consequences (CCPL), and AI infrastructure for physical AI at Syntheta. Everything below is public, runnable, and benchmarked.
Open-source physics-informed digital twin for turbojet engines: health monitoring, prognostics, explainable AI, and interactive 3D visualization.
Repository ↗End-to-end AO pipeline: SH-WFS centroiding with 127× C speedup → CNN/MMSE reconstruction (Strehl 0.9987) → LQG control → SLODAR profiling. BAH 2026 Challenge #9.
Repository ↗Counter-drone swarm defense engine in pure Python/NumPy: IMM-EKF tracking, GRU-based constrained-RL decision layer (CCPL), Hungarian assignment with salvo logic, APN guidance, EW layer. FastAPI gateway, sub-200ms track-to-decision latency.
Project site ↗Fixes credit assignment in safety-critical RL when an action's cost arrives after a delay: delay-corrected Bellman targets, state-conditioned penalties, and SCM-labeled consequence attribution. Open-source on PyPI, 42 passing tests, under review at NeurIPS 2026.
Repository ↗ | Read on Medium ↗
I'm an independent researcher and builder working across reinforcement learning, robotics, and applied machine learning. I focus on problems where learning systems have to interact with the physical world, from constrained decision-making and sensor fusion to digital twins, multi-agent systems, and embedded hardware.
I build end-to-end: physics models, training code, experiments, software, dashboards, and hardware. I write my own code, run my own benchmarks, and publish what I can verify. When something is simulated rather than field-tested, I say so.
I'm interested in building systems that can actually work outside the notebook.
Open-source physics-informed digital twin platform for turbojet engines: engine health monitoring, prognostics, explainable AI, and interactive 3D visualization. Combines thermodynamic modeling with machine learning for real-time diagnostics.
End-to-end adaptive optics pipeline: Shack-Hartmann WFS centroiding with 127× C speedup → CNN/MMSE wavefront reconstruction (Strehl 0.9987) → LQG closed-loop control → SLODAR turbulence profiling. BAH 2026 Challenge #9.
A counter-drone swarm defense engine built in pure Python and NumPy - no PyTorch, no cloud. IMM-EKF tracker with four motion models (CV/CA/CT/Singer) and Joseph-form covariance updates, a custom GRU-based constrained-RL decision layer (CCPL), Hungarian-assignment multi-interceptor coordination with salvo logic, 3D augmented proportional navigation guidance, a physics-first fragment kill model, and an electronic-warfare layer (GPS denial, datalink jamming, DEW turret). Served via a FastAPI REST+WebSocket gateway with sub-200ms track-to-decision latency, plus PyQt6 field operator and DRDO evaluator interfaces.
The problem: in safety-critical RL, a bad action's cost often arrives seconds later - so the agent blames the wrong decision and never learns what actually caused the harm. Standard methods assign every observed cost to the most recent action, which breaks credit assignment exactly when it matters most.
What I built: CCPL, a constrained-RL framework that models the delay explicitly. It corrects the Bellman target with a history-dependent effective factor $\gamma_{\mathrm{eff}}(h) = \sum_{\tau=0}^{K} p(\tau \mid h)\,\gamma^{\tau}$, replaces the single global penalty with a state-conditioned multiplier $\lambda(s)$, attributes consequences to actions via a Consequence Net trained on labels from a controlled structural causal model, and keeps reward and constraint critics separate so updating λ never changes either TD target. Installable from PyPI (`pip install ccpl-rl`), with Gymnasium adapters and 42 passing tests. Submitted to NeurIPS 2026.
Local multi-agent AI system with 6 specialized agents, 12 tools, ChromaDB vector memory, and Ollama LLM backend. Fully offline capable with tool use, memory, and inter-agent coordination.
Production-grade bird vocalization classifier that identifies species and decodes communication meaning (alarm, mating, territorial...) using EfficientNet-B0 on mel-spectrograms, served via FastAPI and containerized with Docker.
Deep learning system that classifies infant cry reasons (hunger, pain, tiredness...) and age group from raw audio using EfficientNet-B0 on mel-spectrograms. Designed as a assistive tool for caregivers and healthcare workers.
A GRU-based RL agent trained entirely on intrinsic reward — zero task reward — using Random Network Distillation, where the agent is rewarded for states its own predictor network can't yet explain: $r^{i}_t = \|\hat{f}(s_t) - f(s_t)\|^2$ against a fixed random target network $f$. Stable across 14,700+ updates with no divergence, using the same delay-corrected Bellman operator from CCPL for value learning. Claims about world-model accuracy, concept formation, and exploration were checked against a battery of pre-registered, falsifiable verification scripts rather than reported post hoc — including a forward world model that predicted true transitions with 29–32× lower error than shuffled controls. A meta-learning "self-editing" layer (SEAL, per Zweiger et al., NeurIPS 2025) used rejection-sampling RL to adapt training hyperparameters online, improving world-model accuracy by 23.7% and exploration coverage by 8.7%.
A wearable personal-safety device combining GPS location tracking with GSM-based SOS alerting to pre-set emergency contacts. Recognized in the India Book of Records; published as first-author research in IJNRD (see Awards).
A reinforcement learning paper addressing three problems in how constrained RL methods compute safety penalties: unknown causal delay between action and consequence, conflation of an agent's own effects with consequences already in motion, and non-stationary Bellman targets when the penalty weight is updated. Proposes a delay-corrected Bellman operator with a proven contraction property, a state-conditioned penalty shown to dominate scalar penalties, and a causal-contribution estimator (an Interventional Consequence Net) trained on ground-truth labels from a structural causal model.
Open to research collaboration, mentorship conversations, and feedback on any of the work above. I reply to email.