
Senior Engineering Manager, Agentic and Generative AI Benchmarking and Evaluations
Real job — pulled straight from ServiceNow’s careers page · Verified July 24, 2026 · No reposts.
Job description
ServiceNow is hiring a Senior Engineering Manager, Agentic and Generative AI Benchmarking and Evaluations — a full-time, based in Santa Clara, CA role ($201k–$352k). Apply directly on ServiceNow's careers page below.
Senior Engineering Manager, Agentic & Generative AI Benchmarking and Evaluations
Location: Santa Clara, California, us
Company Description
It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for everyone—freeing people from busywork so they could focus on meaningful work. Today, is the AI control tower for business reinvention. Our AI platform brings together any AI, any data, and any workflow— helping 85% of the Fortune 500® work smarter, faster, and better. We're building an AI-native culture where technology and talent are unstoppable together. And we're just getting started.
Join us to put AI to work for people.
Job Description
The Advanced Technology Group (ATG) at is a customer-focused innovation group building intelligent software and smart user experiences using existing and latest advanced technologies to enable end-to-end, industry-leading work experiences for customers. We are a group of researchers, applied scientists, engineers, and product managers with a dual mission. We build and evolve the AI platform, and partner with teams to build products and end-to-end AI-powered work experiences. In equal measure, we lay the foundations, research, experiment, and de-risk AI technologies that unlock new work experiences in the future.
Job Description
We are seeking an exceptional, data-driven Senior Engineering Manager, Agentic & GenAI Benchmarking and Evaluations to establish and lead AI evaluation practices for both and our customers. As shifts enterprise workflows from simple generation to complex, autonomous agents, ensuring system reliability, safety, and accuracy is paramount.
In this role, you will lead a specialized team of AI evaluation engineers and data scientists. Your team will build the infrastructure, rigorous validation frameworks, and benchmarks that quantify the performance of Now Assist agentic workflows across multi-step orchestration, tool-calling, and enterprise-grounded reasoning. You will bridge the gap between frontier AI research and hard production metrics, directly impacting the trust and adoption of autonomous workflows for millions of enterprise users.
What You Get To Do In This Role
- Build the Evaluation Infrastructure: Design, own, and scale automated testing and evaluation harnesses (unit evals, integration evals, and production drift monitors) to measure agent quality and eliminate regressions.
- Define Enterprise AI Benchmarks: Create standard, repeatable evaluation frameworks tailored to complex business workflows—assessing multi-agent orchestration, intent routing, multi-step planning loops, and long-term memory accuracy.
- Validate Grounding & RAG Pipelines: Partner with search and data fabric teams to systematically evaluate Retrieval-Augmented Generation (RAG) pipelines, hybrid search, and semantic re-ranking systems.
- Model Selection Optimization: Rigorously benchmark frontier LLMs (e.g., OpenAI, Anthropic, Google, and proprietary models) to evaluate trade-offs across execution capabilities, latency, context-window efficiency, and inference costs.
- Lead a High-Performing Team: Recruit, mentor, and foster an AI-native engineering team, driving engineering best practices, prompt-infrastructure stability, and production-grade rigor.
- Cross-Functional Leadership: Collaborate with Core Product, Machine Learning Platforms, and Engineering leads to translate baseline performance statistics into actionable product improvements and model fine-tuning targets.
Qualifications
To be successful in this role you have:
- 8+ years of professional software engineering or machine learning experience, including 3+ years managing or technically leading high-performing AI/ML teams.
- Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving. This may include using AI-powered tools, automating workflows, analyzing AI-driven insights, or exploring AI's potential impact on the function or industry.
- Strong foundational knowledge of frontier AI SDKs and deep experience deploying or testing agentic/probabilistic software architectures (multi-agent orchestration, tool execution, and probabilistic feature deployment).
- Demonstrated experience implementing rigorous AI metrics (e.g., ROUGE, BLEU, G-Eval, LLM-as-a-judge patterns, and custom deterministic evaluation code) at an enterprise scale.
- Proficiency in Python and familiarity with data analytics infrastructures (SQL, Pandas, NumPy) alongside standard MLOps tracking platforms.
- Experience with complex knowledge infrastructure, SaaS platform architectures, or relational datasets (e.g., Knowledge Graphs, CMDBs).
- Ability to translate deeply technical evaluation data into executive-level risk assessments, ROI summaries, and strategic roadmap recommendations.
- Bachelor’s or higher degree in Computer Science, Data Science, Machine Learning, or a highly quantitative field (Master's or Ph.D. is a plus).
For positions in this location, we offer a base pay of $201,300 - $352,300, plus equity (when applicable), variable/incentive compensation and benefits. Sales positions generally offer a competitive On Target Earnings (OTE) incentive compensation structure. Please note that the base pay shown is a guideline, and individual total compensation will vary based on factors such as qualifications, skill level, competencies, and work location. We also offer health plans, including flexible spending accounts, a 401(k) Plan with company match, ESPP, matching donations, a flexible time away plan and family leave programs. Compensation is based on the geographic location in which the role is located and is subject to change based on work location.
Additional Information
Work Personas
We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to employees depending on the nature of their work and their assigned work location. Learn more here. To determine eligibility for a work persona, may confirm the distance between your primary residence and the closest office using a third-party service.
Equal Opportunity Employer
is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, national origin, age, disability, gender identity, veteran status, or any other category protected by law. In addition, all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements.
Accommodations
We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process, or are unable to use this online application and need an alternative method to apply, please contact globaltalentss@.com for assistance.
Export Control Regulations
For positions requiring access to controlled technology subject to export control regulations, including the U.S. Export Administration Regulations (EAR), may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon obtaining any export license or other approval that may be required by relevant export control authorities.
From Fortune. ©2026 Fortune Media IP Limited. All rights reserved. Used under license.
Get Engineering Manager jobs like this→
New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.
Email me new jobsSimilar jobs




Frequently asked questions
What is the salary for Senior Engineering Manager, Agentic and Generative AI Benchmarking and Evaluations at ServiceNow?
The estimated salary range for Senior Engineering Manager, Agentic and Generative AI Benchmarking and Evaluations at ServiceNow is $201,000 - $352,000 USD per year.
What skills are required for Senior Engineering Manager, Agentic and Generative AI Benchmarking and Evaluations at ServiceNow?
The required skills for Senior Engineering Manager, Agentic and Generative AI Benchmarking and Evaluations at ServiceNow include: Python, SQL, Pandas, NumPy, MLOps, LLM, RAG, API.
What is the seniority level for Senior Engineering Manager, Agentic and Generative AI Benchmarking and Evaluations at ServiceNow?
Senior Engineering Manager, Agentic and Generative AI Benchmarking and Evaluations at ServiceNow is a Senior / Manager level position.
How do I apply for Senior Engineering Manager, Agentic and Generative AI Benchmarking and Evaluations at ServiceNow?
You can view the full description and apply for Senior Engineering Manager, Agentic and Generative AI Benchmarking and Evaluations at ServiceNow on EchoJobs: https://echojobs.io/job/servicenow-senior-engineering-manager-agentic-generative-ai-benchmarking-and-evaluations-u2kfj.