whoami

Explaining what happens [inside] stochastic systems.

I graduated from the Indian Institute of Technology Jodhpur, majoring in CS. I'm the co-founder (technical research) at Safety and Alignment Research, India, where I work on AI safety and machine hermeneutics, inclduing interpretability.

cd ~/research for a more academia-friendly version.

Quick Highlights

all projects >

use multi-layer inputs (and outputs) for your natural language autoencoders.

Mechanistically Interpreting Compression in VLMs

ICML mech interp workshop 2026

circuit and crosscoder feature analysis to interpret compression in VLMs. we also introduce VLMSafe-420, a dataset of multimodal counterfactuals for ai safety.

Fantastic Biases and Where to Find Them

mechanistic interpretability

tracing racial and gender bias in gpt-2 and gpt-neo using probes, activation patching, head ablations, and circuits.

studying alignment fine-tuning methods such as PPO, SimPO, ORPO, GRPO, KTO, and DPO, through probes, SAEs, and crosscoders.

Now

researching
Applied Interpretability, Meta-Models.
organizing
NewInML Workshop @ NeurIPS 2026.
reading
"The Alignment Problem"; "Breakneck: China's Quest to Engineer the Future".

Skills

languages
Python, C, C++ | SQL | HTML, CSS, NextJS.
frameworks & tools
PyTorch, Django, AWS, Git, Bash.
deep learning
HuggingFace, W&B, ollama, unsloth, multi-GPU training, TRL.
ai engineering
LangGraph | Context Engineering | MCP | PineCone, Qdrant | Agent Memory.