I build the data pipelines and the AI systems on top of them: large-scale web scraping and data extraction, RAG with knowledge graphs and source-level citations, AI agents and MCP servers, entity-resolved healthcare datasets, and backtesting tooling for trading strategies.
Top Rated Plus · 100% Job Success · 2,800+ hours · finance, healthcare, real estate, e-commerce. Team of 5 engineers behind me.
Most AI projects stall at the same wall: the team wants an agent or a RAG system, but the data lives in scattered PDFs, broken APIs, half-scraped websites and legacy databases. I do both halves of that problem, so the AI build does not pause for three months while someone "figures out the data".
WEB SCRAPING & DATA EXTRACTION
Production scrapers for Amazon, Etsy, Mercari and niche marketplaces, real-estate portals, gaming and sports platforms — Playwright / Scrapy with anti-bot rotation, proxies, scheduling, deduplication and clean datasets into PostgreSQL, Google Sheets or your API. A real-estate data lake with daily refresh across 12 sources; an RSS / news aggregator; 200K+ legacy financial documents parsed into PostgreSQL with full lineage. Most of my long-term clients started with a scraping job.
RAG, KNOWLEDGE GRAPHS, AI AGENTS & MCP SERVERS
Document-intelligence systems where every answer is traceable: semantic index (pgvector) plus a Neo4j knowledge graph of extracted facts, so answers surface documents that share no vocabulary with the question. Every passage carries file, paragraph and the chain of entities behind it. Connectors for SharePoint, DataSite and Outlook, background workers, multi-tenant isolation, on-prem option. AI agents (Claude, OpenAI) and MCP servers that expose your data and APIs to any MCP client. AI assistant for an electronics distributor (BOM cross-referencing, 84% catalog coverage); AI ops assistant for affiliate networks (Affise / Keitaro / Voluum / Tune).
HEALTHCARE DATA AGGREGATION & ENTITY RESOLUTION
Medical Data Graph for a MedTech due-diligence platform: 7.4M US providers (NPPES), 9.3M Medicare procedure records, 17M Open Payments transactions, 2.3M facility affiliations, 194K clinical trials with 133K investigator roles resolved to named physicians, 1.88M PubMed publications with 3.25M author links — linked by a custom matching engine with confidence scoring (95–100% accuracy on high-confidence tiers) and served through a 9-endpoint FastAPI / PostgreSQL API in 2–3 seconds. Services: NPI matching and HCP data validation, physician / KOL target lists, market analysis, the graph deployed in your cloud. Public data only, no PHI. The same aggregation and record-linkage engine works for companies, contracts, people, properties.
Andrii H. earns an estimated $7.9k/mo. That's 5.4× the typical freelancer and more than 99.79% of everyone we track.
QUANT BACKTESTING & STRATEGY MIGRATION
Migrated a legacy trading-signal system (VB.NET / EasyLanguage) to Python with signal-by-signal validation against the original; built corporate-action-adjusted market-data pipelines, backtesting with standard risk metrics, and parameter-optimization tooling with out-of-sample checks. I port EasyLanguage, AFL, MQL, ThinkScript and Excel strategies to Python and prove the port before anyone trades it.
HOW I WORK: week 1 scope and edge cases locked; weeks 2–3 something running on your data, not slides; then iteration with evals, guardrails and observability. Agents are engineering, not prompts.
STACK: Python, FastAPI, Claude / Anthropic API, MCP, OpenAI, Azure OpenAI, LangChain / LangGraph, pgvector, Qdrant, Neo4j, PostgreSQL, Redis, Airflow, Playwright / Scrapy, pandas / Polars, AWS, Azure, Docker.
Send me what you are trying to build — I reply within 24 hours on whether I am the right fit.