FriendliAI logo

Forward Deployed Engineer - AI Inference

FriendliAI

On-site
San Francisco, CA
Full-time
Mid Level
3+ yrs
Salary not listedPosted 38m ago

Real job — pulled straight from FriendliAI’s careers page · Verified October 8, 2026 · No reposts.

Job description

FriendliAI is hiring a Forward Deployed Engineer - AI Inference — a full-time, based in San Francisco, CA role. Apply directly on FriendliAI's careers page below.

Forward Deployed Engineer - AI Inference

Department: Engineering

Location: San Francisco

Employment Type: FullTime

About the job

FriendliAI is seeking a Forward Deployed Engineer to assist enterprises in deploying, scaling, and operating generative and agentic AI workloads on FriendliAI infrastructure. You will work directly with customers to solve and implement production-grade applications using our products, such as Serverless Endpoints, Dedicated Endpoints, or Container.

Friendli Container is our service that allows customers to download our inference engine as Docker images and deploy it in their chosen environment, such as private clouds or on-premises. Our Friendli Container can be adopted directly to AWS EKS clusters using our EKS add-on product.

You will work directly on our customers’ projects, collaborating with their engineering teams to solve AI inference challenges like scaling, orchestration, and monitoring. This is a hands-on, customer-embedded role. If you have worked in DevOps, platform engineering, or SRE for AI applications, this is your ideal position.

Key Responsibilities

  • Design and implement large-scale deployment architectures for LLM and multimodal inference

  • Deploy and manage containerized workloads across Kubernetes clusters

  • Diagnose production issues, such as performance bottlenecks, and implement temporary fixes as needed

  • Collaborate with customers’ DevOps teams to integrate FriendliAI’s infrastructure into their CI/CD workflows

  • Develop scripts, Helm charts, and Terraform modules that simplify repeated deployments

  • Contribute field insights to shape our platform reliability, observability, and scaling strategies

  • Lead workshops, technical sessions, or webinars to help customers master infrastructure best practices.

Qualifications

  • 3+ years of experience in cloud infrastructure, DevOps, or reliability engineering

  • Bachelor’s or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent

  • Proficiency with Kubernetes, Docker, Terraform, and Helm

  • Strong foundation in distributed systems, networking, and performance tuning

  • Experience with GPU-based computing and generative AI model serving workloads

  • Strong technical background in backend systems or AI tooling

  • Experience operating workloads on AWS, GCP, or OCI

  • Excellent problem-solving and debugging skills in real-world environments

Preferred Experience

  • Experience deploying large models (LLMs, diffusion models) on GPUs or clusters

  • Familiarity with inference frameworks (Triton, vLLM, TensorRT, DeepSpeed-Inference)

  • Familiarity with observability stacks (Prometheus, Grafana, Loki, ELK, OTEL)

  • Understanding of networking security and compliance frameworks (e.g., SOC 2)

  • Experience supporting on-prem or hybrid-cloud deployments

Benefits

  • A front-row seat to the generative AI infrastructure revolution

  • Competitive compensation and benefits package

  • Daily lunch and dinner provided; unlimited snacks and beverages

  • Health check-up and top-tier hardware support

  • Flexible working hours and a highly collaborative environment

About us

FriendliAI is the fastest inference cloud for agents, built to run frontier open-weight models in production at scale. It delivers up to 7x faster output token speed, up to 90% lower inference costs, and 99.99% uptime across the most demanding agent workloads — long-context inference, real-time streaming, and accurate tool calling.

We are a small, fast-moving team doing work that matters at one of the most exciting moments in the history of technology. With our world-class inference stack, we are building the platform teams can actually rely on.

Get Software Engineer jobs like this→

New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.

Email me new jobs
Sia Partners logo

Staff Software Engineer (Python)

Mumbai, MH
✓ From careers page· 1d ago
Sia Partners logo

Engineering Manager, Product

Mumbai, MH
✓ From careers page· 1d ago
Sia Partners logo

Senior Backend Engineer (Python)

Mumbai, Maharashtra
✓ From careers page· 1d ago
Sia Partners logo

Staff DevOps Engineer

Mumbai, Maharashtra
✓ From careers page· 1d ago

Frequently asked questions

What skills are required for Forward Deployed Engineer - AI Inference at FriendliAI?

The required skills for Forward Deployed Engineer - AI Inference at FriendliAI include: Kubernetes, Docker, Terraform, Helm, AWS, GCP, Oracle Cloud, Generative AI, LLM, Prometheus, Grafana, Elasticsearch, OpenTelemetry, SOC 2, CI/CD.

What is the seniority level for Forward Deployed Engineer - AI Inference at FriendliAI?

Forward Deployed Engineer - AI Inference at FriendliAI is a Mid Level level position.

How do I apply for Forward Deployed Engineer - AI Inference at FriendliAI?

You can view the full description and apply for Forward Deployed Engineer - AI Inference at FriendliAI on EchoJobs: https://echojobs.io/job/friendliai-forward-deployed-engineer-ai-inference-89zbw.