Geisinger logo

Tech Lead Data Scientist, AI Evaluation & Monitoring

Geisinger

Hybrid
Full-time
Senior
Staff
Salary not listedPosted 2mo ago

Real job — pulled straight from Geisinger’s careers page · Verified June 11, 2026 · No reposts.

Job description

Geisinger is hiring a Tech Lead Data Scientist, AI Evaluation & Monitoring — a full-time, remote role. Apply directly on Geisinger's careers page below.

Tech Lead Data Scientist, AI Evaluation & Monitoring

Location: Danville, Pennsylvania, United States

Employment Type: Full time

What You Will Own: 

  • The technical evaluation methodology applied to AI programs across the enterprise, pre-production validation, production monitoring, and ongoing optimization 
  • Hands-on guidance to program teams as they design validation studies, equity audits, monitoring plans, and escalation playbooks for their AI systems 
  • Instrumentation of production monitoring: translating program-specific failure modes into concrete, measurable metrics 
  • The evaluation toolkit: LLM-as-Judge frameworks, golden sets, simulation harnesses, experimental study designs, drift detection, subgroup fairness analysis 
  • Reusable evaluation playbooks and templates that let each new program move faster than the last 
  • Technical direction, design review, and mentorship for a team of data analysts supporting the evaluation function 

What You Will Not Own: 

  • People management, HR administration, or formal performance evaluations for the analyst team (those sit with the analysts' line manager; the Tech Lead provides technical input) 
  • Program-level product strategy or go/no-go decisions 
  • Final clinical validation judgment on whether a given AI is safe for a given clinical use 
  • The software infrastructure behind the evaluation and monitoring tooling (built by the AI Platform team — the Tech Lead defines what's measured and how; Platform builds the backend) 

Shape of the Work:

This is a role that lives at three altitudes at once: 

With program teams (hands-on advisory). Partner with program owners early, before evaluations are designed, to shape study approach, sample size, stratification, gold-standard definition, and decision thresholds. Translate ambiguous failure modes into concrete, defensible evaluation designs. Coach teams through the technical work so that what arrives at governance review is rigorous, not performative. 

With the evaluation toolkit (hands-on build). Design and operate the reusable assets that let evaluation scale: LLM-as-Judge rubrics and calibration methods, golden sets, simulation harnesses, A/B and shadow-mode study templates, subgroup fairness analyses, and drift monitors. Keep a pragmatic eye on what actually works in a clinical environment versus what works in a paper. 

With the analyst team (technical leadership). Set technical direction, assign work across active evaluations, review analysis code and study designs, and raise the technical bar. Mentor analysts on methodology, statistical rigor, and the domain knowledge that makes evaluation credible. Grow them from execution into independent evaluation design. 

Methods You'll Use: 

  • Experimental and quasi-experimental design for production AI systems 
  • LLM and generative AI evaluation: golden sets, judge-based evaluation, hallucination and grounding checks 
  • Fairness and equity evaluation across patient and stakeholder subgroups 
  • Production monitoring design: drift detection, performance decay, adoption, and outcome metrics 
  • Causal inference methods appropriate to healthcare settings where full RCTs are often impractical 
  • Simulation and adversarial testing for pre-production stress testing 
  • Python, SQL, modern ML and evaluation tooling, cloud-native data platforms 

Work is typically performed in an office or remote environment. Accountable for satisfying all job specific obligations and complying with all organization policies and procedures. The specific statements in this profile are not intended to be all-inclusive. They represent typical elements considered necessary to successfully perform the job.

*Relevant experience may be a combination of related work experience and degree obtained (Master's Degree = 2 years; PHD = 4 years ).

Get Data Scientist jobs like this

New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.

Email me new jobs
Westpac New Zealand logo

Platform Engineer

Auckland, NZ
✓ From careers page· 39m ago
Westpac New Zealand logo

Senior Quality Engineer, Automation Tester

Hapori Pōneke
✓ From careers page· 40m ago
Westpac New Zealand logo

Senior Data Engineer

Auckland, NZ
✓ From careers page· 41m ago
Westpac New Zealand logo

Senior Data Engineer

Auckland, NZ
✓ From careers page· 41m ago

Frequently asked questions

Is Tech Lead Data Scientist, AI Evaluation & Monitoring at Geisinger a remote job?

Yes, Tech Lead Data Scientist, AI Evaluation & Monitoring at Geisinger is a remote position. Candidates in Danville, PA may be preferred.

What skills are required for Tech Lead Data Scientist, AI Evaluation & Monitoring at Geisinger?

The required skills for Tech Lead Data Scientist, AI Evaluation & Monitoring at Geisinger include: Python, SQL, LLM, AI, Machine Learning, Deep Learning.

What is the seniority level for Tech Lead Data Scientist, AI Evaluation & Monitoring at Geisinger?

Tech Lead Data Scientist, AI Evaluation & Monitoring at Geisinger is a Senior / Staff / Principal level position.

How do I apply for Tech Lead Data Scientist, AI Evaluation & Monitoring at Geisinger?

You can view the full description and apply for Tech Lead Data Scientist, AI Evaluation & Monitoring at Geisinger on EchoJobs: https://echojobs.io/job/geisinger-tech-lead-data-scientist-ai-evaluation-monitoring-vanoj.