Cantina logo

Media Software Engineer, Speech

Cantina

On-site
Sunnyvale, CA
Full-time
Senior
Staff
3+ yrs
$180k–$270kPosted 19h ago

Real job — pulled straight from Cantina’s careers page · Verified August 23, 2026 · No reposts.

Job description

Cantina is hiring a Media Software Engineer, Speech — a full-time, based in Sunnyvale, CA role ($180k–$270k). Apply directly on Cantina's careers page below.

Media Software Engineer, Speech (Senior-Staff Levels)

Department: Engineering

Location: Sunnyvale, San Francisco

Employment Type: FullTime

About Cantina:

Cantina Labs is a social AI company, developing a suite of advanced real-time models that push the boundaries of expression, personality, and realism. We bring characters to life, transforming how people tell stories, connect, and create. We build and power ecosystems. Cantina, our flagship social AI platform, is just the beginning.

If you're excited about the potential AI has to shape human creativity and social interactions, join us in building the future!

About the Role:

The Media Team at Cantina is building the real-time infrastructure powering live conversations between people and AI characters. Our goal is simple to express, but challenging to make real: enabling fast, natural, and truly conversational interaction with diverse and creative characters.

We’re looking for a Software Engineer to help improve the speech, audio, and media systems at the heart of the Cantina experience.

This team’s responsibilities cover low-level media processing pipelines, integration with internal and external models for speech recognition and synthesis, and globally distributed, WebRTC-based infrastructure supporting real-time voice and video interactions across iOS, Android, and web.

If you’re excited by high-performance C++, real-time systems, speech technologies, and building the future of conversational AI, we’d love to talk.

What You’ll Do:

  • Improve the real-time speech and media systems powering live AI conversations.

  • Reduce latency and optimize responsiveness across audio streaming and speech pipelines.

  • Build tools to enable the creation of AI models to support speech processing.

  • Develop new voice and video capabilities that enable more immersive interactions between users and AI bots.

  • Improve and extend our custom WebRTC infrastructure across iOS, Android, and web.

What You’ll Bring:

Minimum qualifications:

  • BS or MS in Computer Science, Computer Engineering, or a related field; or equivalent experience.

  • 3+ years of experience working as a software engineer.

  • Excellent communications skills.

  • Demonstrated ability to work independently to drive projects from requirements to completion.

  • Experience with C or C++ in a professional context.

  • Grounding in computer science fundamentals, including memory management, high-performance data structures, and concurrent / multithreaded systems.

  • Exposure to system programming concepts, including network protocol design, asynchronous I/O, and distributed system architectures..

  • Object-oriented development and design skills.

  • Interest in solving subtle and challenging engineering problems.

Preferred qualifications:

  • Previous experience with WebRTC, streaming protocols, or other media-adjacent technologies.

  • Familiarity with media processing techniques..

  • Experience creating backend server infrastructure.

  • Experience developing software for iOS or Android.

  • Familiarity with building services using Node.js or Go.

  • Familiarity with artificial intelligence and machine learning techniques, particularly in relation to speech recognition and synthesis.

Compensation:

The anticipated annual base salary range for this role is between $180,000-$270,000. When determining compensation, a number of factors will be considered, including skills, experience, job scope, location, and competitive compensation market data.

Benefits:

  • Competitive salary and generous company equity

  • Medical, dental, and vision insurance – 99.99% of premiums covered by Cantina

  • 42 days of paid time off, including:

    • 15 PTO days

    • 10 sick days

    • 15 company holidays

    • 2 floating holidays

  • Generous parental leave & fertility support

  • 401(k) retirement savings plan

  • Lifestyle spending account – $500/month to use however you’d like

  • Complimentary lunch and snacks for in-office employees

  • One Medical membership, and more!

Get Media Software Engineer, Speech jobs like this

New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.

Email me new jobs
Pipe Care Group logo

Mechanical Engineer

Amman, Amman, Jordan
✓ From careers page· 39m ago
Pipe Care Group logo

Mechanical Engineer

Cairo, Cairo, Egypt
✓ From careers page· 39m ago
Socomec logo

Product Development Engineer

$110k–$130kBrampton, ON
✓ From careers page· 55m ago
Terracon logo

Senior Landscape Architectural Consultant

Tempe, AZ
✓ From careers page· 1h ago

Frequently asked questions

What is the salary for Media Software Engineer, Speech at Cantina?

The estimated salary range for Media Software Engineer, Speech at Cantina is $180,000 - $270,000 USD per year.

What skills are required for Media Software Engineer, Speech at Cantina?

The required skills for Media Software Engineer, Speech at Cantina include: C++, C, Node.js, Go, iOS, Android.

What is the seniority level for Media Software Engineer, Speech at Cantina?

Media Software Engineer, Speech at Cantina is a Senior / Staff level position.

How do I apply for Media Software Engineer, Speech at Cantina?

You can view the full description and apply for Media Software Engineer, Speech at Cantina on EchoJobs: https://echojobs.io/job/cantina-media-software-engineer-speech-senior-staff-levels-uw6r4.