Clera logo

Data Scientist, Agent Evaluations & Quality

Clera

On-site
Palo Alto, CA
Full-time
Senior
4+ yrs
Salary not listedPosted 5d ago

Real job — pulled straight from Clera’s careers page · Verified August 8, 2026 · No reposts.

Job description

Clera is hiring a Data Scientist, Agent Evaluations & Quality — a full-time, based in Palo Alto, CA role. Apply directly on Clera's careers page below.

Data Scientist — Agent Evaluations & Quality

Department: Engineering

Location: Palo Alto

Employment Type: FullTime

About the Role

We're a small, fast-moving AI productivity startup (~25 people) building an autonomous AI executive assistant that operates across email, calendars, meetings, and business software. This role owns the measurement system that determines whether our AI agent is genuinely improving in ambiguous, real-world environments.

You'll partner closely with AI Agent Capabilities engineers to produce the evidence that drives product decisions, model choices, and release quality — turning hard questions about agent behavior into rigorous, actionable answers.

Visa sponsorship is not available for this role.

What You'll Do

  • Architect and maintain automated evaluation pipelines that measure agent quality across product surfaces.

  • Translate product capabilities into explicit pass, partial-pass, and failure criteria for complex multi-step tasks.

  • Build representative gold datasets and regression suites covering real workflows, edge cases, and adversarial scenarios.

  • Define and track metrics including task success, tool-selection accuracy, instruction adherence, factual consistency, latency, cost, and reliability.

  • Design deterministic and model-based graders, calibrate LLM-as-a-judge systems, and monitor grader agreement.

  • Compare models, prompts, and implementations using rigorous offline experiments and production evidence.

  • Analyze traces and production outcomes to identify root causes and build a practical failure taxonomy.

  • Convert production failures into regression cases and continuously close gaps in evaluation coverage.

  • Build dashboards and release-quality signals that make results actionable for engineering, product, and leadership.

  • Partner with capability engineers to recommend improvements and verify that fixes raise quality without introducing unacceptable regressions.

What We're Looking For

Required experience (dealbreakers):

  • 4+ years in Applied Data Science or Machine Learning roles, with a focus on building and delivering evaluation systems, automated data pipelines, or production ML infrastructure.

  • Demonstrated experience designing and implementing automated evaluation frameworks, success criteria, and regression suites for complex AI/ML or agentic systems.

  • Production-grade proficiency in Python and SQL, with hands-on experience building and maintaining automated analytical pipelines on large datasets.

Core requirements:

  • Experience applying statistical and experimental methods — significance testing, variance analysis, sampling — to evaluate non-deterministic AI/ML systems.

  • Experience developing labeled datasets, annotation guidelines, and quality-control processes for ground-truth data in dynamic product environments.

  • Deep understanding of LLM agent behaviors including tool use, multi-step execution, retrieval, and practical failure modes.

  • Ability to analyze model traces, tool calls, and outputs to identify root causes of failures across model, prompt, tool, and data layers.

  • Experience using production telemetry and observability data to monitor system quality, build dashboards, and analyze real-world user outcomes.

Nice to have:

  • Prior hands-on experience with LLM-as-a-judge systems, model-based grading, or AI benchmarking platforms.

  • Experience shipping or operating production ML products, agentic systems, or customer-facing consumer software.

  • Experience reviewing and adapting public research benchmarks or academic evaluation methodologies to real-world product problems.

Key traits we value: strong product orientation (prioritizing metrics tied to real user outcomes over convenient proxies), high ownership, analytical rigor, and comfort driving ambiguous quality questions from design through to product decisions.

Location

This role is based in Palo Alto, CA. Visa sponsorship is not available.

Get Data Scientist jobs like this

New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.

Email me new jobs
Ledgebrook logo

Head of Forward Deployed Engineering (Remote)

$200k–$250kRemote · US-eligible
✓ From careers page· 14m ago
CERN logo

CERN

New

Software Developer

Geneva, GENEVA, ch
✓ From careers page· 34m ago
CERN logo

CERN

New

Digital ASIC Designer

Geneva, CH
✓ From careers page· 34m ago
CERN logo

CERN

New

Full-Stack Software Engineer

Geneva, GENEVA
✓ From careers page· 34m ago

Frequently asked questions

What skills are required for Data Scientist, Agent Evaluations & Quality at Clera?

The required skills for Data Scientist, Agent Evaluations & Quality at Clera include: Python, SQL, Machine Learning, Data Science, LLM, AI.

What is the seniority level for Data Scientist, Agent Evaluations & Quality at Clera?

Data Scientist, Agent Evaluations & Quality at Clera is a Senior level position.

How do I apply for Data Scientist, Agent Evaluations & Quality at Clera?

You can view the full description and apply for Data Scientist, Agent Evaluations & Quality at Clera on EchoJobs: https://echojobs.io/job/clera-data-scientist-agent-evaluations-quality-ndazm.