Senior Machine Learning Engineer, AI Platform & Agentic Apps
Building agentic AI systems with LLMs, evals, and safety guardrails — directly aligned with vibe coding for agent tooling and AI-driven developer workflows.
About the Role
Build and productionize Robinhood's agent platform and agentic applications, focusing on safe, trustworthy agent behavior through trajectory-level evaluation and action guardrails. Lead technical direction, mentor engineers, and deliver platform primitives that enable internal and customer-facing agents to take real actions at scale in a regulated financial environment.
Job Description
Role
Senior/Staff Machine Learning Engineer on the AI Platform & Agentic Apps team responsible for designing and building the harness that powers Robinhood’s agentic AI. The role focuses on making agents trustworthy at scale via evaluation systems and action-level guardrails, and requires partnering with product, infrastructure, and ML engineers to move ideas from prototype to production.
Key Responsibilities
- Design and build the core agent harness: orchestration, tool integrations, context and memory management.
- Ship agentic applications end-to-end, taking ambiguous problems to production agents that take real actions for employees or customers.
- Build trajectory-level evaluation systems (tool-call correctness, planning and recovery, multi-step task completion) with simulation environments and synthetic task generation.
- Architect action guardrails as platform primitives: least-privilege tool scoping, permission models, human-approval gates, sandboxing, rollback, and budget/step limits.
- Productize evals and guardrails so other teams adopt them: SDKs, CI regression gates for prompt/model/tool changes, continuous red-teaming, and production tracing that feeds back into evals and guardrail models.
- Set technical quality through architecture and code reviews, mentorship, and data-driven release decisions.
Requirements
- 10+ years of experience as a Machine Learning Engineer or ML-focused software engineer; Master’s degree in CS or related field or equivalent experience.
- Strong Python and distributed-systems fundamentals and a track record of shipping LLM-powered systems to production at scale.
- Hands-on experience building agentic systems end-to-end: tool use, orchestration, context management, and multi-step planning in production.
- Deep expertise in evaluating agents: trajectory-level evals, tool-call scoring, and simulation environments.
- Demonstrated experience designing action-level guardrails: permission/tool-scoping, approval gates, blast-radius controls, and sandboxing.
- Rigor in evaluation methodology: golden datasets, rubric/LLM-as-judge grading, statistical significance with small N, offline-to-online metric correlation, and eval data versioning/contamination control.
- Proven ability to build platforms (eval, safety, or agent tooling) that other teams adopt and sound judgment on build vs buy decisions.
Compensation & Location
- Base pay ranges by zone: Zone 1 (Menlo Park, CA; New York, NY; Bellevue, WA; Washington, DC): $255000 — $300000 USD; Zone 2 (Denver, CO; Westlake, TX; Chicago, IL): $225000 — $264000 USD; Zone 3 (Lake Mary, FL; Clearwater, FL; Gainesville, FL): $199000 — $234000 USD.
- Role is based in the Menlo Park, CA office with in-person attendance expected at least 3 days per week.
Benefits
- Performance-driven compensation, bonus opportunities, and equity ownership.
- 401(k) matching.
- 100% paid health insurance for employees and 90% coverage for dependents.
- Access to the Robinhood Employee Fund and top AI tools plus continuous AI skill-building.
- Lifestyle wallet flexible spending account for wellness and learning.
- Employer-paid life & disability insurance, fertility benefits, and mental health benefits.
- Paid time off including company holidays, PTO, sick time, and parental leave.
- Exceptional office experience with catered meals and events.