I build Playwright test infrastructure and evaluation systems for AI-enabled products. Recent work includes scaling Jenkins execution from 8 to 32 parallel workers, architecting a 32-worker Playwright pipeline on AWS EKS, implementing CI/CD release gates, and building full-stack LLM applications with TypeScript and Python.
I’m a Senior SDET with 10+ years in software quality engineering and 5+ years designing test infrastructure. My primary stack includes Playwright, TypeScript, Python, AWS, Kubernetes, Selenium Grid, REST APIs, PostgreSQL, Kafka, and CI/CD.
Recent results include:
• Architected an AWS EKS Playwright pipeline running 32 concurrent workers
• Scaled Jenkins test execution from 8 to 32 parallel workers
• Built CI/CD quality gates for smoke, regression, API, and discovery suites
• Maintained a 32-session Selenium Grid while migrating coverage to Playwright
• Integrated automated results with Zephyr Scale, Slack, and CloudFront
• Built full-stack AI product capabilities with Next.js, FastAPI, PostgreSQL, and multi-provider LLM integrations using OpenRouter
I can help with:
• Playwright or Selenium automation architecture
• QA infrastructure, CI/CD pipelines, and release gates
• Automated UI, API, integration, and data-pipeline testing
• LLM, RAG, chatbot, and AI-agent evaluation
• Prompt and response regression suites
• Deterministic AI test harnesses
• Testing and reviewing AI-generated code
• Full-stack TypeScript and Python development
• Stabilizing flaky, slow, or difficult-to-maintain test suites
For AI systems, I focus on measurable behavior, not simply generating more tests with an LLM. This can include curated evaluation datasets, deterministic assertions, rubric-based scoring, provider and model comparisons, hallucination and failure-path testing, latency and cost checks, and regression gates that run in CI.
I use AI extensively to increase delivery speed while relying on tests, reviews, observability, and measurable acceptance criteria to validate its output.
If your team needs reliable automation, scalable test execution, or an engineer who can both test and build the product, send me your current stack and main release bottleneck. I can propose a practical first milestone.
Andrew H. earns an estimated $7.6k/mo. That's 5.2× the typical freelancer and more than 99.77% of everyone we track.