QA AI Automation Engineer
Explicitly requires vibe coding skills — building and evaluating multi-agent AI workflows, prompt engineering, and AI-assisted code generation.
About the Role
Build and own QA and evaluation systems for LLM-powered assistant and API surfaces in a regulated wealth-management platform. Design and operate multi-agent CI workflows that detect, triage, and optionally fix AI-generated defects while enforcing code quality for agent-authored code.
Job Description
Role
This is a hands-on engineering role to evaluate, test, and gate non-deterministic LLM-driven systems within a regulated wealth management platform. You will own the AI evaluation framework, wire evaluations into CI, and build multi-agent workflows that detect and remediate defects automatically while scaling code quality controls for agent-authored code.
Key Responsibilities
- Own and extend the AI evaluation framework for chat and API surfaces: correctness, grounding/citation fidelity, retrieval quality, refusal/escalation behavior, tone, latency, and cost per interaction.
- Curate and version golden datasets and adversarial test sets representing real advisor workflows.
- Design LLM-as-judge and rubric-based scoring for non-deterministic outputs and validate judges against human-labeled data.
- Integrate evaluations into CI to gate changes in prompts, models, retrieval, and tools; instrument production traces and close the loop from live failures back to evals.
- Design and operate a multi-agent quality step in CI/CD that executes suites, triages failures, writes defects with reproduction detail, proposes or applies fixes, and re-runs verification.
- Define agent roles, hand-off logic, guardrails, and trust thresholds for auto-applying fixes versus routing to humans; make cost-aware decisions about LLM use vs deterministic automation.
- Define and implement automated gates for agent-authored code: coverage and mutation tests, static analysis, security and dependency scanning, architectural conformance, and review checklists.
- Identify failure modes specific to agent-generated code and build detection for them; partner with architects and delivery teams to maintain velocity without sacrificing quality.
- Maintain and modernize existing automation (API, UI, integration) and mentor engineers on quality practices.
Requirements
Required
- 5+ years in software quality, test automation, or SDET work with ownership of automation architecture.
- Production experience with LLM-powered systems: agent patterns, prompt engineering, tool/function calling, and orchestration.
- Hands-on experience with major model APIs (OpenAI, Anthropic, Azure OpenAI) and AI-assisted development tooling (Claude Code, Copilot, Codex or equivalent).
- Strong programming in at least one: Python, TypeScript/JavaScript, Java, Kotlin, or C#; comfortable reading others.
- Experience building and interpreting evaluation criteria for systems without a single correct answer.
- Deep CI/CD fluency: pipeline integration, gating, monitoring, and logging.
- API testing experience and validating third-party integrations.
- Clear written communication to explain quality risk and root causes across stakeholders.
Preferred
- Experience with eval and observability tooling such as Langfuse, Promptfoo, OpenTelemetry, New Relic, or equivalents.
- Familiarity with RAG fundamentals: embeddings, chunking, vector search, and retrieval evaluation.
- Experience with workflow orchestration tooling (n8n or similar).
- Exposure to Azure cloud services and the .NET ecosystem.
- Experience with BDD/Selenium or comparable UI automation at scale.
- Background in financial services or other regulated domains and experience mentoring or leading distributed QA engineers.
First 90 Days
- Days 1–30: Learn the assistant, API layer, and current eval/automation estate; ship a baseline eval suite for a high-traffic chat workflow with CI-visible results.
- Days 31–60: Stand up the first multi-agent defect-detection step on a single service; establish trust threshold for auto-applied fixes and publish metrics.
- Days 61–90: Propose and begin rolling out agent-authored code quality standards with automated gates across delivery teams.