Benchling logo

Software Engineer, Model Evaluation and Improvement

Benchling

On-site
San Francisco, CA
Full-time
Mid Level
2+ yrs
$136k–$167kPosted 1h ago

Real job — pulled straight from Benchling’s careers page · Verified August 26, 2026 · No reposts.

Job description

Benchling is hiring a Software Engineer, Model Evaluation and Improvement — a full-time, based in San Francisco, CA role ($136k–$167k). Apply directly on Benchling's careers page below.

Software Engineer, Model Evaluation and Improvement

Department: Engineering

Location: San Francisco, CA

Compensation: $136,435 – $166,754 • Offers Equity

Employment Type: FullTime

We are rebuilding biotech for the AI era.

When a breakthrough is delayed, the world waits. Getting a molecule from discovery to patients, or a crop from lab to field, involves thousands of slow, manual, disconnected steps. AI has the potential to change this, compressing decades of R&D work into years. But that only happens when clean, structured scientific data and AI are built into how science gets done.

Benchling is the AI platform for biotech R&D. Scientists use Benchling to design experiments, capture structured data, and run AI agents and models directly in their workflows. Over 200,000 scientists around the world trust Benchling to power their most important work, from academic labs to Sanofi, Moderna, and more than half of the world's top 50 biopharma.

We’re building an AI scientist for our customers. We can’t do that if we haven’t built the muscle ourselves. AI fluency is the foundation we build on; it's core to how we work, and we're committed to helping every new hire integrate it into their day-to-day. As part of our interview process, you'll complete a brief AI-focused exercise or discussion so we can understand how you think about and use AI to drive impact in your role. Feel free to reference any tools, platforms, or workflows you use today.

Role Overview

We’re a team focused on making frontier AI models better at science. LLMs know an extraordinary amount of biology, but there’s still a large gap in reasoning for the real-world problems scientists face every day. We recently published some of our work here.

You’ll build the datasets, evaluations, and systems that help close that gap. You’ll work with scientists to turn complex scientific work into rigorous tasks that models can learn from and be evaluated against. You’ll partner with leading AI labs to understand where models fail and how to improve them.

This is an early and rapidly evolving area. You’ll work at the intersection of software engineering, biology, and frontier AI: finding tasks that are challenging for LLMs and valuable to scientists, designing evaluations that capture real scientific judgment, and building systems to create these tasks at scale.

 

RESPONSIBILITIES

  • Build datasets for evaluating and improving frontier models, turning complex scientific data into high-quality tasks and environments for LLMs.

  • Analyze model failure modes, running experiments across frontier models to understand where they struggle and identify opportunities for improvement.

  • Build scalable data infrastructure, creating pipelines that curate, transform, and validate large volumes of scientific data into tasks for model evaluation and improvement.

  • Collaborate with frontier AI labs, helping develop and evaluate new approaches for improving models on challenging scientific tasks.

  • Work closely with scientists, translating expert judgment into problems and evaluation criteria that can reliably distinguish strong model behavior.

QUALIFICATIONS

  • 2+ years at the intersection of biology and AI, with experience evaluating and improving scientific models or LLMs for biological applications.

  • Experience building with LLMs, with an intuition for where current models excel, where they struggle, and how to design systems around their capabilities.

  • Curiosity and excitement about frontier AI, with a desire to understand and push the capabilities of rapidly improving models.

  • Comfort working on ambiguous problems, where the playing field is rapidly shifting and the right technical approaches are still being discovered.

  • Collaborative mindset, able to work closely with engineers, scientists, and external research partners.

  • Desire to work in a fast-paced environment, where priorities can shift and rapid experimentation is encouraged.

HOW WE WORK

This is an in-person team in San Francisco built around collaborating in the office in a fast-paced environment. We’re in the office Monday through Friday.

#LI-KW1

Benchling welcomes everyone.

We believe diversity enriches our team so we hire people with a wide range of identities, backgrounds, and experiences.

We are an equal opportunity employer. That means we don’t discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status. We also consider for employment qualified applicants with arrest and conviction records, consistent with applicable federal, state and local law, including but not limited to the San Francisco Fair Chance Ordinance.

Get Software Engineer jobs like this

New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.

Email me new jobs
Dispensed logo

Senior Analytics Engineer

Australia
✓ From careers page· 28m ago
Jack Henry logo

Senior AI Engineer

Remote · US-eligible
✓ From careers page· 28m ago
CloudFactory logo

Principal Software Engineer

Remote · Canada-eligible
✓ From careers page· 29m ago
CloudFactory logo

Lead Software Engineer

Costa Rica
✓ From careers page· 29m ago

Frequently asked questions

What is the salary for Software Engineer, Model Evaluation and Improvement at Benchling ?

The estimated salary range for Software Engineer, Model Evaluation and Improvement at Benchling is $136,000 - $167,000 USD per year.

What skills are required for Software Engineer, Model Evaluation and Improvement at Benchling ?

The required skills for Software Engineer, Model Evaluation and Improvement at Benchling include: Python, Machine Learning, AI, LLM.

What is the seniority level for Software Engineer, Model Evaluation and Improvement at Benchling ?

Software Engineer, Model Evaluation and Improvement at Benchling is a Mid Level level position.

How do I apply for Software Engineer, Model Evaluation and Improvement at Benchling ?

You can view the full description and apply for Software Engineer, Model Evaluation and Improvement at Benchling on EchoJobs: https://echojobs.io/job/benchling-software-engineer-model-evaluation-and-improvement-xaqdq.