AI researcher and Principal AI Engineer. Indian, based in Dubai. My work sits where post-training, agents, evaluation, and inference meet: training open-weight models to do specific jobs well, building the systems that run them in production, and measuring whether they actually got better. I contribute upstream to the projects that stack depends on.
20 merged upstream across 4 repos. Updated 2026-10-08.
| Repo | Merged | Summary |
|---|---|---|
| unslothai/unsloth | 12 merged | Fixes rough edges in the chat and studio UI: stopped or queued prompts stay put, snapshot-path models keep their context and compare settings, and uninstall, eject, and streaming paths behave more predictably. |
| google/adk-python | 5 merged | Fixes edge cases across agent config, auth, MCP, and LiteLlm so ADK fails early on bad settings (candidate_count) and behaves correctly in less-common setups (Windows VCS paths, public OIDC clients, non-Gemma tool-result roles, MCP grounding metadata). |
| twentyhq/twenty | 2 merged | Fixes edge cases around the percent sign: percentage field values entered with a trailing "%" now save correctly, and array containsIlike filters treat "%" as a SQL wildcard as intended. |
| e2b-dev/E2B | 1 merged | Sandbox creation now checks the options you pass before it asks for an API key, so a bad config fails fast with a clear error instead of an unrelated auth complaint. |
Fine-tuned and shipped Qwen3.6-27B with QLoRA/SFT, then on-policy context distillation from a 24.8K-token expert-rule corpus. Merged adapters into a standalone bf16 checkpoint and served it on H100s with vLLM at 73K context.
Built the self-improving loop around it: production generations become critic-ranked Best-of-N preference pairs and broken-to-repaired training examples for DPO and RLVR/GRPO, with deterministic verifiable rewards and frozen held-out gates deciding whether a candidate model is promoted or rejected. On a locked 30-task by 3-seed harness, rule distillation cut mean defects 54% (0.79 to 0.36 per output) and raised strict zero-defect passes from 17% to 53%.
DeepField: an independent research program on whether the early signatures of significant events can be detected ahead of consensus.
Earlier, a CNN for object amplification in IBM Watson PowerAI Vision.
CiaraAI. Stateful agent platform running autonomous multi-step workflows across CRM, lead routing, scheduling, payments, voice, WhatsApp, and chat. 200K+ interactions, 97% automation, sub-2s responses, about 70% lower cost than the stack it replaced.
Claws. Visual multi-agent orchestration on the OpenClaw runtime for manager-worker-reviewer teams. Workflow definitions compile to executable agent configs with concurrent execution, tool integrations, runtime monitoring, and failure recovery.
Raven. Open-source AI meeting copilot: system audio plus mic capture, WebRTC AEC3 echo cancellation, real-time transcription, live assistance. 400+ stars.
Evaluation and observability. A 559-rule learning-science evaluation framework catching about 92% of defects before human QC, and root-cause instrumentation across 500K+ production AI failures, turned into regression signals.
Stack. Python, TypeScript, PyTorch, Unsloth, vLLM, LangGraph, PostgreSQL, Kubernetes, AWS, GCP. Claude Certified Architect (Foundations).




