whoami
Explaining what happens [inside] stochastic systems.
I graduated from the Indian Institute of Technology Jodhpur, majoring in CS. I'm the co-founder (technical research) at Safety and Alignment Research, India, where I work on AI safety and machine hermeneutics, inclduing interpretability.
cd ~/research for a more academia-friendly version.
Quick Highlights
all projects >use multi-layer inputs (and outputs) for your natural language autoencoders.
circuit and crosscoder feature analysis to interpret compression in VLMs. we also introduce VLMSafe-420, a dataset of multimodal counterfactuals for ai safety.
tracing racial and gender bias in gpt-2 and gpt-neo using probes, activation patching, head ablations, and circuits.
studying alignment fine-tuning methods such as PPO, SimPO, ORPO, GRPO, KTO, and DPO, through probes, SAEs, and crosscoders.
Now
- researching
- Applied Interpretability, Meta-Models.
- organizing
- NewInML Workshop @ NeurIPS 2026.
- reading
- "The Alignment Problem"; "Breakneck: China's Quest to Engineer the Future".
Skills
- languages
- Python, C, C++ | SQL | HTML, CSS, NextJS.
- frameworks & tools
- PyTorch, Django, AWS, Git, Bash.
- deep learning
- HuggingFace, W&B, ollama, unsloth, multi-GPU training, TRL.
- ai engineering
- LangGraph | Context Engineering | MCP | PineCone, Qdrant | Agent Memory.