Open to work: freelance & full-time

Hey, I'm Jay Thakkar.

I build LLM systems that survive production: MCP infrastructure, voice AI agents, and RAG pipelines that stay fast, cheap and honest under real traffic. 3+ years shipping for enterprise contact centres, fintech and SaaS teams. Google Developer Expert for AI & Cloud.

Everywhere as @akajammythakkar
Jay Thakkar
🎙️ Voice AI
🔌 MCP
📚 RAG
  • 0%Query latency cut
    4s → 300ms via MCP
  • 0%LLM compute spend
    reduced
  • 0%Hallucinations cut
    with prompt guardrails
  • 0Developers trained
    APAC · MENAT · SSA
Freelance & consulting

What I can build for you

Scoped engagements delivered end to end, covering architecture, implementation, evals and handover. Five enterprise clients shipped, 95% satisfaction, integration timelines cut by 40%.

🧠

LLM systems, end to end

From a fuzzy idea to a deployed product: model selection, orchestration across GPT / Claude / Gemini / LLaMA, prompt architecture, evals, and the boring reliability work that decides whether it ships.

  • Multi-LLM routing
  • Chain-of-Thought
  • Evals
  • FastAPI
🎙️

Voice AI agents

Contact-centre grade voice agents: intent design, prompt tuning, barge-in and latency budgets, system-level guardrails. Improved response reliability by 15% and cut hallucinations 87.5% on live deployments.

  • Contact centre
  • ServiceNow
  • Guardrails
  • Latency tuning
🔌

MCP integrations

Model Context Protocol servers that connect your models to real systems: 10+ data sources, VM-based auth handling 500+ secure API transactions a day, zero downtime. Query time down 92.5%.

  • MCP servers
  • Tool design
  • Auth & secrets
  • CLI bridges
📚

RAG pipelines

Ingestion, chunking, hybrid retrieval and re-ranking that actually improves answers. I've taken retrieval accuracy from 30% to 90% and lifted generation quality 40% on client corpora.

  • Pinecone / Weaviate
  • LlamaIndex
  • Docling
  • Re-ranking
📉

AI cost optimisation

Your bill is usually a design problem. Caching, model tiering, context discipline and retrieval hygiene, for a 60% inference cost reduction without giving up quality.

  • Spend audit
  • Model tiering
  • Caching
  • Observability
🎓

Advisory, workshops & talks

Google Developer Expert for AI & Cloud. Team enablement, architecture reviews, and hands-on workshops, using the same material that's reached 30,000+ developers across three regions.

  • Architecture review
  • Team training
  • Keynotes
  • GDE

How engagements work

Discovery sprint: one to two weeks, ending with an architecture and a costed plan. Build: fixed-scope delivery with weekly demos. Retainer: ongoing tuning, evals and on-call for the system once it's live.

Book a call →
Experience

Where I've built

Impulsive Web → Technology 9 Labs

Jun 2025 - Present

Full-time · Hybrid · Impulsive Web acquired by Technology 9 Labs, 2025

Resident Entrepreneur, Technology / AI EngineerDec 2025 - Present
  • Promoted from AI Engineer to Resident Entrepreneur on acquisition, and now own AI infrastructure strategy and enterprise client delivery.
  • Architected MCP infrastructure across 10+ data sources, cutting query response time 92.5% (4s → 300ms); VM-based auth processing 500+ secure API transactions daily with zero downtime.
  • Tuned voice AI prompts with Chain-of-Thought and few-shot techniques, giving 15% better LLM response reliability across enterprise contact-centre deployments for 3CLogic.
  • Shipped system-level prompt guardrails cutting hallucinations 87.5%, and a CLI tool bridging backend APIs with Claude and Codex that took solution cycles from 15 weeks to 2.
  • Full-cycle AI integration for 5+ enterprise clients via custom MCP pipelines: 40% faster data onboarding, 95% client satisfaction.
AI Engineer, Impulsive WebJun 2025 - Nov 2025
  • Scoped and architected end-to-end MCP-based AI pipelines for enterprises across fintech, logistics and SaaS, cutting integration timelines by 40%.
  • Led cross-functional sprints shipping production-ready AI modules on schedule, NPS above 95% across engagements.

Forsyt.ai: Revenue Operating System

  • Built the AI Revenue Intelligence engine (Python, React, PostgreSQL, Gemini) ingesting CRM, email and calendar data to surface risk and auto-generate action bundles, lifting deal adherence 15%.
  • Action Flow Engine for one-click AI-drafted email and CRM updates, saving 10+ hours per rep per week.
  • Confidence Scoring model on sentiment and engagement velocity, improving prediction of deal outcomes by 25%.

Hoodo: product startup

  • Launched an intelligent calendar platform automating 100+ meetings per participant.
  • AI meeting-notes system with full audio capture and structured note generation, zero content loss.
  • MCP
  • Voice AI
  • Gemini
  • FastAPI
  • PostgreSQL
  • GCP

SCAAI (Symbiosis Centre for Applied AI)

Jul 2023 - Present

Part-time · 2 yrs 7 mos

R&D Mentor, Gen AIJun 2024 - Present · Pune, Maharashtra · Hybrid
Research InternJul 2023 - Jul 2024 · Remote
  • Generative AI
  • Research mentoring
  • R&D

Researcher

Jan 2025 - May 2025

Truxt.ai · Full-time · Remote

Binoloop

Jan 2024 - Jan 2025

Canada · Remote · 1 yr 1 mo

Machine Learning AssociateOct 2024 - Jan 2025 · Full-time
AI/ML Research AnalystJan 2024 - Oct 2024 · Internship
  • Promoted from research analyst to associate within the same tenure for RAG pipeline and multi-LLM orchestration work.
  • Built and deployed RAG pipelines with advanced text extraction, giving 40% better generation quality at 60% lower inference cost.
  • Orchestrated solutions across Gemini, GPT, Claude and LLaMA, a 30% uplift in accuracy and efficiency across 8+ client engagements.
  • RAG
  • LangChain
  • Vector stores
  • Cost optimisation

AI Intern

Nov 2023 - Feb 2024

Otobit Private Limited · Full-time · Surat, Gujarat · On-site

  • Python
  • Decision Sciences
Earlier roles, 2020 to 2023
  1. Student CoordinatorPPSU Data Science Club · Full-time · Surat, Gujarat
    Aug 2022 - Sep 2023
  2. Project ManagerTriclone Technologies Private Limited · Full-time · Surat, Gujarat · On-site
    Dec 2022 - May 2023
  3. Teaching InternP P Savani University · Part-time · Surat, Gujarat · On-site
    Oct 2022 - Apr 2023
  4. Business Intelligence SpecialistPGP Glass Pvt. Ltd. · Internship · Vadodara, Gujarat
    Jun 2022 - Jul 2022
  5. Data Scientist → Data Analyst → Software Developer TOPS Technologies Pvt. Ltd · Internship · Surat, Gujarat · Data Scientist (Oct 2021 - Aug 2022), Data Analyst (Oct 2021 - Apr 2022), Software Developer (Jun 2021 - Oct 2021)
    Jun 2021 - Aug 2022
  6. Software DeveloperTOPS Technologies Pvt. Ltd · Internship · Surat, Gujarat · Java developer who built a Java-based website
    Nov 2020 - Mar 2021

Toolkit

Languages

  • Python
  • SQL
  • Go
  • Node.js

ML & LLM

  • PyTorch
  • TensorFlow
  • JAX
  • Scikit-Learn
  • LangChain
  • LlamaIndex
  • NLTK
  • Unstructured IO
  • Docling

AI platforms

  • Anthropic Claude
  • OpenAI GPT
  • Google Gemini
  • Meta LLaMA
  • Vertex AI
  • Hugging Face
  • FastAPI

Data

  • Pinecone
  • Weaviate
  • PostgreSQL
  • Supabase
  • Vector stores

Cloud & DevOps

  • GCP
  • AWS
  • Docker
  • Git
  • MLflow
Selected work

Projects

Open source, all of it on github.com/akajammythakkar.

TPU Sprint Q1 2026

vForge

Benchmarks LLM fine-tuning on TPU against GPU end to end: a chat-driven dataset builder, LoRA fine-tunes in JAX and Keras 3, then throughput and cost compared with vLLM.

  • JAX
  • Keras 3
  • Cloud TPU
  • vLLM
View repo
★ 12 4 forks 24K views

AI Hiring Agent

Five specialist agents under one conversational orchestrator, scoring resumes and validating GitHub profiles live before returning a hiring verdict.

  • Google ADK
  • Gemini 3
  • Multi-agent
View repo
★ 22 Most starred 50K views

RAG with Gemini

Retrieval-augmented PDF analysis on Gemini, worth a 5% lift in question-answering accuracy over the baseline. The most-read thing I have published.

  • Python
  • RAG
  • Gemini
  • GCP
View repo
Codelab

Multi-Agent Coding Assistant

Planner, coder and explainer agents collaborating in real time, each streaming its own live UI panel into a Next.js app over the 2026 agent protocol stack.

  • A2A
  • AG-UI
  • Google ADK
  • Next.js
View repo
Fully offline

Local RAG with Gemma

Retrieval-augmented answers where no document leaves the machine: Gemma 3 through Ollama, a FastAPI backend and a numpy vector store with no external database.

  • Gemma 3
  • Ollama
  • FastAPI
  • Local-first
View repo
★ 5 3 forks

Task Prioritisation Agent

Ranks a task list by urgency, importance, effort and company OKRs with Vertex AI Gemini, served as both a REST API and a small containerised web app.

  • Vertex AI
  • Gemini
  • FastAPI
  • Docker
View repo
Community

Google Developer Expert for AI & Cloud

Since March 2024 I've been teaching AI/ML in the open through Google-backed programmes, community talks and hands-on workshops that have reached 30,000+ developers.

  • Drove AI literacy across APAC, MENAT and SSA through training programmes and technical workshops.
  • Grew reach from APAC-only in 2024 to three regions inside 12 months, tripling geographic coverage and speaking engagements year over year.
Invite me to speak →

Education

M.Tech in Artificial Intelligence & Machine Learning Symbiosis Institute of Technology Aug 2024 - May 2026
B.Tech in Computer Science Engineering (AI & ML) P. P. Savani University Oct 2020 - Jun 2024
Based in Surat, Gujarat, India Working with teams across IST, CET, EST and PST
Research

Peer-reviewed publications

  1. Evaluating Large Language Models for Knowledge-aware Question and Answering

    International Journal on Smart Sensing and Intelligent Systems

  2. Optimizing Energy Management in Smart Grids: A Hybrid Approach to Load Forecasting

    PACIS 2025

  3. A Pragmatic Approach on Adoption of EDA to Make Intelligent Business Decisions

    International Journal of Wireless Network Security

Open to work: freelance & full-time

Let's build something that ships.

Have an LLM idea stuck at the demo stage, a voice agent that hallucinates, or a bill that keeps climbing? Tell me what you're working on. I reply to everything.