I build production AI agents and LLM systems. 50+ projects across fintech, healthcare, media, legal tech, climate tech, and research: multi-agent systems, agent harnesses, evaluation systems for agentic workflows, Claude Code, Claude Agent SDK, OpenAI Codex, Hermes, Openclaw, MCP servers, RAG at scale, fine-tuning, and document AI.
7+ years shipping ML and AI in production. $700K+ delivered. Expert-Vetted (top 1%). 100% Job Success Score.
I’m not new to ML. Before the LLM wave I spent years building classical ML and NLP systems in Python: search engines, recommendation engines, entity extraction, knowledge graphs, and predictive models. That foundation means I know when a problem needs an LLM and when it doesn’t. I’ve also cofounded two startups, so I think about business outcomes, not just model accuracy.
Recent work:
N1 Healthcare: Led the Report Generation team on a medical AI platform. Built the agent harness that dispatches 13+ clinical report workflows across Claude Agent SDK, OpenAI Agents SDK, and Agno through a LiteLLM gateway, with a planner coordinating 20+ specialty agents. Report cost went from roughly $80 to $15-25 per run. Also worked on the upstream medical record parsing pipeline.
Newsweek (4+ years): 15+ AI applications and 6 Azure Function Apps used daily by non-technical editors. Cosmos DB vector search with Reciprocal Rank Fusion hybrid ranking on 3072-dim embeddings, MCP servers exposing agent tools to Copilot Studio, automated article-to-video pipeline at 95%+ success. A live 30-day window measured 497 editorial hours saved at 15.8% AI adoption. Trained editorial staff to run the systems without developer support.
Stratifi: Built the AI layer on a Django + Postgres backend. Dual-LLM PDF extractor (GPT-4o + Gemini) at 95%+ accuracy on 50+ page brokerage statements, 10-30 positions/second, with a 24-document eval harness gating prompt changes. Multi-agent financial chatbot routing across 4 market data sources, plus a 4-layer portfolio optimization agent and a Lambda support agent with tool calling. Elasticsearch hybrid search over millions of securities.
CGIAR: Research impact assessment processing 9,166+ records with 6 fine-tuned GPT-4o-mini variants, GPU-accelerated PDF layout detection with Detectron2, 3-tier evidence extraction at 50 concurrent operations.
Adoro: 7-stage AI automation pipeline for marketing, turning a week-long manual process into under an hour, processing 100K+ items per run with 3-pass LLM extraction (initial, reflection, self-reflection).
Nikhil B. earns an estimated $7.9k/mo. That's 5.3× the typical freelancer and more than 99.79% of everyone we track.
ClimateX: Web crawling and location intelligence pipeline with Flair NER for entity extraction, dual geocoding APIs, and proximity-based deduplication.
Accelchain: Smart contract vulnerability detection covering 37 SWC vulnerability types with vector-DB retrieval over 33,488 known patterns, RAG plus few-shot plus Chain of Thought for explainable remediation.
SynMax: NLP pipeline processing 400,000+ historical records with LongT5 fine-tuned at 16K context, 16-category document classification, cross-document entity linking across court records, deeds, and financial documents.
Epitome (cofounded): HR Tech B2B2C platform with custom Word2Vec on millions of LinkedIn profiles, ArangoDB knowledge graph modeling O*NET across 20+ edge collections, entity resolution cascade (exact, fuzzy, embedding, LLM), Elasticsearch candidate search, 40+ REST endpoints.
Core capabilities: AI Agents and Orchestration, Agent Harnesses, Agent Evals, Multi-Agent Systems, Claude Agent SDK, MCP, LLM Fine-Tuning and Post-Training, RAG and Hybrid Retrieval, Document AI and OCR, Production ML and Forecasting, NLP and Knowledge Graphs, Data Pipelines and ETL, Cloud Infrastructure, and API Development.