Clera logo

Data Scientist, Agent Evaluations & Quality

Clera

On-site
Palo Alto, CA
Full-time
Senior
5+ yrs
Salary not listedPosted 1h ago

Real job — pulled straight from Clera’s careers page · Verified August 23, 2026 · No reposts.

Job description

Clera is hiring a Data Scientist, Agent Evaluations & Quality — a full-time, based in Palo Alto, CA role. Apply directly on Clera's careers page below.

Data Scientist — Agent Evaluations & Quality

Department: Engineering

Location: Palo Alto

Employment Type: FullTime

About the Role

We're an early-stage AI company building autonomous agents that handle real work — email, calendar, browser, business software, and more. We're hiring a Data Scientist focused on Agent Evaluations & Quality to measure, understand, and continuously improve the quality of our agent capabilities.

Your mission is to translate ambiguous product behavior into measurable definitions of success, build representative evaluation datasets, design reliable graders and metrics, analyze failures, and create the feedback loops that guide engineering and product decisions. This is applied data science at the intersection of evaluation design, statistics, experimentation, production Python, and deep understanding of how LLM agents behave in real products.

This is a full-time, on-site role based in Palo Alto, CA. Visa sponsorship is not available.

What You'll Do

  • Architect and maintain automated evaluation pipelines that measure agent quality across capabilities and product surfaces.

  • Translate capabilities into explicit success criteria — including pass, partial-pass, and failure definitions for complex multi-step tasks.

  • Build representative gold datasets and regression suites covering common workflows, ambiguous requests, long-tail behavior, edge cases, and adversarial scenarios.

  • Define and track metrics such as task success, partial completion, tool-selection accuracy, tool-use correctness, instruction adherence, factual consistency, user corrections, latency, cost, and reliability.

  • Design deterministic graders, model-based graders, and human-review processes; calibrate LLM-as-a-judge systems and measure false positives, false negatives, variance, and grader agreement.

  • Analyze traces, tool calls, model outputs, user context, and production outcomes to identify root causes and build a useful failure taxonomy.

  • Compare models, prompts, tools, and capability implementations using rigorous offline experiments and production evidence.

  • Build dashboards, reports, and release-quality signals that make evaluation results understandable and actionable for engineering, product, and leadership.

  • Partner with capability engineers to recommend improvements and verify that fixes raise quality without unacceptable regressions in cost, latency, or reliability.

What We're Looking For

Required — Dealbreakers:

  • 5+ years of experience in data science, machine learning, or analytics roles building or delivering evaluation systems, metrics frameworks, or quality measurement solutions for production systems.

  • Demonstrated experience designing and implementing evaluation frameworks, metrics, and grading systems for ML/AI systems in production.

  • Production-quality Python and SQL proficiency with the ability to build automated data pipelines and analysis code at scale.

Required Skills & Experience:

  • Experience designing evaluation methodologies: success criteria definition, dataset construction, metric selection, and distinguishing useful benchmarks from misleading ones.

  • Statistical and experimental design knowledge: sampling, variance, uncertainty quantification, bias detection, confounding variables, and significance testing for non-deterministic systems.

  • Experience with ground-truth data development: labeling guideline design, annotation quality control, ambiguity resolution, and dataset maintenance as product behavior evolves.

  • Working knowledge of LLM behavior — including model-based graders, tool use, retrieval systems, multi-step execution, partial completion, and practical failure modes.

  • Analytical debugging ability: connecting quantitative patterns to individual system traces and identifying failure origins across model, prompt, context, tools, data, and application logic.

  • Experience building dashboards, reports, and communicating evaluation results, methodology, uncertainty, and trade-offs to both technical and non-technical stakeholders.

Nice to Have:

  • Experience with LLM-as-a-judge systems, calibration, and measurement of grader agreement, false positives, and false negatives.

  • Prior work on evaluation or benchmarking platforms for AI systems.

  • Experience with agentic systems, multi-step task execution, or tool-use evaluation.

  • Experience working on customer-facing consumer software or production ML systems with real-world user outcomes.

Location & Work Arrangement

  • On-site in Palo Alto, CA

  • Visa sponsorship is not available

Get Data Scientist jobs like this

New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.

Email me new jobs
Exadel logo

Senior Software Engineer (Python, Java, Golang)

Brazil
✓ From careers page· 30m ago
OAK'S LAB logo

Tech Lead

Serbia
✓ From careers page· 31m ago
OAK'S LAB logo

Tech Lead

Portugal
✓ From careers page· 31m ago
OAK'S LAB logo

Fullstack Engineer

Turkey
✓ From careers page· 32m ago

Frequently asked questions

What skills are required for Data Scientist, Agent Evaluations & Quality at Clera?

The required skills for Data Scientist, Agent Evaluations & Quality at Clera include: Python, SQL, Machine Learning, Data Science, LLM, QA.

What is the seniority level for Data Scientist, Agent Evaluations & Quality at Clera?

Data Scientist, Agent Evaluations & Quality at Clera is a Senior level position.

How do I apply for Data Scientist, Agent Evaluations & Quality at Clera?

You can view the full description and apply for Data Scientist, Agent Evaluations & Quality at Clera on EchoJobs: https://echojobs.io/job/clera-data-scientist-agent-evaluations-quality-dwuob.