Software Engineer, Machine Learning Infrastructure - Generative AI
DoorDash USA · San Francisco Bay Area
📍 San Francisco, CA; Sunnyvale, CA; Seattle, WA💰 $137,100via greenhousePosted 2026-07-22
Apply on company site ↗
CareerRiver pulls this listing straight from the employer's hiring system — no recruiter middleman, no reposts. Applying takes you directly to DoorDash USA.
About the Team
DoorDash’s GenAI Platform team sits within Machine Learning Platform and builds the shared infrastructure that helps DoorDash, Wolt, and Deliveroo teams safely bring GenAI-powered products, agents, automation, and personalization to production. Our mission is to increase the velocity of business impact from GenAI. A central pillar of that work is our evaluation platform — the unified evals backbone that lets teams measure, trace, and trust the quality of LLM and agent systems across the company, powering trace/score ingestion, LLM-as-judge workflows, agent simulations, and LLM observability for the tens of millions of daily requests flowing through our LLM Gateway. We also own core platform surfaces including the Agent Gateway, open-weights model serving and batch inference, guardrails, and cost attribution.
About the Role
You will join a small, high-leverage team building production infrastructure for Generative AI at DoorDash, with a primary focus on our evals and LLM observability platform: the systems that let teams evaluate, trace, and continuously improve the quality of LLM and agent products. You’ll work across evaluation frameworks and SDKs, OpenTelemetry-based trace/score ingestion, LLM-as-judge and offline/online eval pipelines, agent simulations, data pipelines, backend services, and observability. This role is ideal for an engineer who enjoys building reliable measurement and quality primitives in a fast-moving technical area where product needs, model capabilities, vendor ecosystems, and evaluation methodologies are evolving quickly.
You’re excited about this opportunity because you will…
Build the infrastructure that helps DoorDash teams move GenAI ideas from prototype to production, increasing the velocity of business impact from AI across the company.
Work on our unified evals platform — evaluation SDKs, OpenTelemetry trace/score ingestion, LLM-as-judge, offline and online eval pipelines, and agent simulations — alongside the LLM Gateway, Agent Gateway, open-weights model serving, guardrails, and cost attribution.
Design scalable systems for evaluation workflows, trace/score ingestion, LLM observability, and agent simulation that power real customer and internal automation use cases
Raise the quality bar for GenAI at DoorDash — giving product teams trustworthy, low-friction ways to measure model and agent quality, catch regressions, and compare across open-weight and closed-source models with observability and cost controls built in.
Build platforms that support rapid experimentation while meeting production standards for latency, scale, monitoring, SLOs, playbooks, and operational excellence.
Partner closely with ML engineers, product engineers, data scientists, and platform teams across DoorDash, Wolt, and Deliveroo to turn emerging GenAI capabilities into durable platform primitives.
Shape the future of DoorDash’s centralized GenAI platform — closing the loop from evaluation and agent observability to agent optimization, where eval signals and traces drive automated evaluation, agent simulation, and post-training techniques (e.g., reward modeling and RLHF/RLVR evaluation) — enabling the next generation of AI-powered products, agents, automation, and personalization.
We’re excited about you because…
B.S., M.S., or PhD. in Computer Science or equivalent
3+ years of industry experience in software engineering
Strong backend engineering fundamentals, especially in Python and distributed systems.
Experience building production services, APIs, data pipelines, or ML infrastructure at scale.
Experience operating systems in production, including observability, debugging, reliability, incident response, and performance/cost optimization.
Hands-on experience with evaluation, LLM observability, or measurement systems for ML/LLM products in production — eval pipelines, tracing/scoring, offline/online quality metrics, or experimentation.
Proficiency in using AI coding tools (e.g., Claude Code, Codex, Cursor) in the full software development lifecycle, including designing, generating code, testing, monitoring and releasing software
Nice To Haves
Depth in evaluation methodology — LLM-as-judge design and calibration, judge/eval drift detection, human-in-the-loop labeling, or eval harness design for agents and multi-step systems
Experience with LLM observability and tracing (e.g., OpenTelemetry, trace/score ingestion) and building instrumentation SDKs
Experience building and deploying AI agents or MCP servers in production, including agent evaluation or simulation
Experience with data pipelines, streaming ingestion, and analytical stores (e.g., SQL, columnar/OLAP) for high-volume telemetry
Experience with LLM gateways, model routing, vendor abstraction, or cost attribution
Experience building developer platforms, internal platforms, or self-serve infrastructure
Experience with Kubernetes, cloud infrastructure (AWS/GCP), or high-throughput batch systems
Experience with RAG, search, vector databases, or open-weights LLM inference and fine-tuning
Notice Regarding Use of AI and Automated Tools: To streamline our hiring process, DoorDash utilizes an automated recruitment tool called Gem.
How it works: Gem assists our recruiting team by evaluating job related qualifications and characteristics in connection with hiring. The tool is designed and used to support - rather than replace - human decision-making; trained personnel make final decisions with meaningful human review and oversight, and DoorDash does not use Gem or other AI-enabled tool in a manner that has the effect of subjecting applicants or employees to discrimination based on any protected characteristic or proxy or for engaging in any protected activity under applicable law.
Data Retention, Privacy & Bias Audit: Data collected during this process is retained in accordance with our Candidate Privacy Policy and applicable state laws. In compliance with
More San Francisco Bay Area jobs
San Francisco Bay Area jobs · Browse all locations