Agents don't fail because the model is bad. They fail because the tools they're handed are unnavigable, the retrieval is unmeasured, and nothing catches a regression before the client does.
I build the layer underneath: production MCP servers (Model Context Protocol), RAG pipelines with measured accuracy, and the eval harnesses that keep both honest.
🏆 Top Rated Plus · $350K+ earned · 8,000+ hours billed
━━━━━━━━━━━━━━━━━━━━━━
RECENT PRODUCTION RESULTS
🔌 Built and shipped a production MCP server (FastMCP, Python) that gives a client's analysts direct agent access to their own domain data — questions that used to need an engineer now get answered in the chat window. It runs in production behind an agent service on AWS Fargate over stdio, and in Claude Desktop and Claude Code. I designed the tool surface for progressive discovery — broad list, then filter and count, then drill down — so models navigate 15K+ records without blowing their context, and built credential-gated tool registration with a read-only-by-design data layer so pointing an agent at live data is safe. 850+ tests, and an eval harness I run across model versions before shipping changes, so a model upgrade can't silently break agent behaviour.
🤝 Wrote the agent that drives it, too — a project-level Claude Code subagent with a curated tool allowlist, in-prompt gates derived from real production failures, and a cost-tier routing matrix that picks the cheapest transport likely to work. Building the tool surface and the agent that consumes it is a different skill from wiring up one API.
🤖 Built a RAG extraction service (FastAPI + Celery + Pinecone, two-tier model routing with a per-model cost estimator) turning messy documents into structured data across ~11K projects from 11 registry sources. Measured on a golden dataset, accuracy went from 33% to 91% F1 on one extraction task and 43% to 84% on another — the difference between a pipeline nobody trusted and one the team runs unattended. Every output traces back to its source document.
🧠 Built a second production RAG system on a medical knowledge graph (FastAPI, Neo4j, MongoDB) with character-level span citations, so a disputed claim takes seconds to check rather than an afternoon. Application-layer tenant isolation with dedicated tests proving no cross-tenant leakage. 2,700+ tests; I wrote roughly two-thirds of the codebase.
🌍 Built and operate a scraping platform covering 20 sources behind enterprise anti-bot protection — Cloudflare-class WAFs and Incapsula, a rotating datacenter proxy pool plus a residential tier for the hardest targets, and fallback transport chains that step up only when they have to. 225K+ documents collected to date. 50+ scheduled pipelines, around 30 of them daily, with per-source error recovery and alerting.
Ross F. earns an estimated $13k/mo. That's 9× the typical freelancer and more than 99.94% of everyone we track.
━━━━━━━━━━━━━━━━━━━━━━
WHAT I DO
✔️ AI integration & agent infrastructure — MCP (Model Context Protocol) servers with FastMCP, tool surfaces designed for how models actually search, Claude Desktop and Claude Code integrations, custom Claude Code subagents, prompt engineering, eval harnesses
✔️ LLM & RAG backends — Anthropic/OpenAI/Gemini APIs, Pinecone, vector search with RRF fusion, structured extraction from messy documents, measured accuracy against golden datasets, cost routing that sends the easy 80% to cheap models
✔️ Web scraping & data extraction — Playwright, Selenium, ZenRows; resilient access to protected sources, proxy management, scheduled fleets via Celery, PDF/Excel/Word extraction, normalization into clean schemas
✔️ API development — FastAPI, Flask, Django; auth, rate limiting, background jobs, clean documentation
✔️ Distributed systems — Celery, RabbitMQ, Redis; retries, idempotency, fault tolerance under real load. Redis caching at 85-95% hit rate, typically 10-50x faster responses
✔️ Production ops — Docker, AWS, PostgreSQL/MongoDB/Neo4j, 110+ zero-downtime migrations on a single project, CI-gated test suites running 2,500-4,900 tests on my largest systems
━━━━━━━━━━━━━━━━━━━━━━
HOW I WORK
✅ I own systems end-to-end: architecture → implementation → deployment → monitoring → handover docs.
✅ I'll tell you when an LLM is the wrong tool — and what to use instead. Cheaper for both of us than finding out in week three.
✅ Failures surface where you'll see them: Prometheus/AlertManager into Slack with per-alert templates, Grafana and Loki for dashboards and logs, scheduled digests and a daily data-feed tripwire. I get paged, not you.
✅ Most of my $350K+ comes from repeat clients and multi-year engagements.
━━━━━━━━━━━━━━━━━━━━━━
📍 UK-based (GMT/BST)
If you need agent tooling that real models can navigate, an LLM pipeline whose accuracy you can actually check, or scraping infrastructure that survives contact with real anti-bot systems — send me a couple of lines about your project and I'll tell you straight away whether I'm the right fit.