LLM systems, end to end
From a fuzzy idea to a deployed product: model selection, orchestration across GPT / Claude / Gemini / LLaMA, prompt architecture, evals, and the boring reliability work that decides whether it ships.
I build LLM systems that survive production: MCP infrastructure, voice AI agents, and RAG pipelines that stay fast, cheap and honest under real traffic. 3+ years shipping for enterprise contact centres, fintech and SaaS teams. Google Developer Expert for AI & Cloud.
Scoped engagements delivered end to end, covering architecture, implementation, evals and handover. Five enterprise clients shipped, 95% satisfaction, integration timelines cut by 40%.
From a fuzzy idea to a deployed product: model selection, orchestration across GPT / Claude / Gemini / LLaMA, prompt architecture, evals, and the boring reliability work that decides whether it ships.
Contact-centre grade voice agents: intent design, prompt tuning, barge-in and latency budgets, system-level guardrails. Improved response reliability by 15% and cut hallucinations 87.5% on live deployments.
Model Context Protocol servers that connect your models to real systems: 10+ data sources, VM-based auth handling 500+ secure API transactions a day, zero downtime. Query time down 92.5%.
Ingestion, chunking, hybrid retrieval and re-ranking that actually improves answers. I've taken retrieval accuracy from 30% to 90% and lifted generation quality 40% on client corpora.
Your bill is usually a design problem. Caching, model tiering, context discipline and retrieval hygiene, for a 60% inference cost reduction without giving up quality.
Google Developer Expert for AI & Cloud. Team enablement, architecture reviews, and hands-on workshops, using the same material that's reached 30,000+ developers across three regions.
Discovery sprint: one to two weeks, ending with an architecture and a costed plan. Build: fixed-scope delivery with weekly demos. Retainer: ongoing tuning, evals and on-call for the system once it's live.
Full-time · Hybrid · Impulsive Web acquired by Technology 9 Labs, 2025
Part-time · 2 yrs 7 mos
Truxt.ai · Full-time · Remote
Canada · Remote · 1 yr 1 mo
Otobit Private Limited · Full-time · Surat, Gujarat · On-site
Benchmarks LLM fine-tuning on TPU against GPU end to end: a chat-driven dataset builder, LoRA fine-tunes in JAX and Keras 3, then throughput and cost compared with vLLM.
View repoFive specialist agents under one conversational orchestrator, scoring resumes and validating GitHub profiles live before returning a hiring verdict.
View repoRetrieval-augmented PDF analysis on Gemini, worth a 5% lift in question-answering accuracy over the baseline. The most-read thing I have published.
View repoPlanner, coder and explainer agents collaborating in real time, each streaming its own live UI panel into a Next.js app over the 2026 agent protocol stack.
View repoRetrieval-augmented answers where no document leaves the machine: Gemma 3 through Ollama, a FastAPI backend and a numpy vector store with no external database.
View repoRanks a task list by urgency, importance, effort and company OKRs with Vertex AI Gemini, served as both a REST API and a small containerised web app.
View repoSince March 2024 I've been teaching AI/ML in the open through Google-backed programmes, community talks and hands-on workshops that have reached 30,000+ developers.
International Journal on Smart Sensing and Intelligent Systems
PACIS 2025
International Journal of Wireless Network Security
Have an LLM idea stuck at the demo stage, a voice agent that hallucinates, or a bill that keeps climbing? Tell me what you're working on. I reply to everything.