
Real job — pulled straight from realm-labs’s careers page · Verified August 24, 2026 · No reposts.
Job description
realm-labs is hiring a AI Red Teaming Intern — a full-time, based in Sunnyvale, CA role. Apply directly on realm-labs's careers page below.
Intern: AI Red Teaming (Fall 2026)
Location: Sunnyvale, CA
Department: AI/ML
Location Type: IN_OFFICE
Employment Type: FULL_TIME
Role Overview
- You will try to break the systems we build and the models we protect, and turn what you find into something the team can act on: a reproducible attack, an evaluation that catches it, and a written account of why it works.
- Expect some mix of: eliciting unsafe behaviour from aligned LLMs and multi-modal models; prompt injection and tool-use abuse against agentic systems; automating attack generation and evaluation rather than hand-crafting one-off prompts; and measuring whether guardrails hold under pressure. Where RealmLabs' interpretability work gives you access to a model's internals, use it.
- We aim for a paper or public technical report out of every internship, plus attacks that stay in our evaluation suite after you leave.
Expected Background: Adversarial ML and Red Teaming
- Hands-on experience attacking or stress-testing models, from any direction: jailbreaks, prompt injection, adversarial examples, data poisoning, model extraction, or evaluating safety and moderation systems.
- Able to read a paper and implement its attack.
- (nice to have) Offensive security background outside ML: CTFs, vulnerability research, penetration testing.
- (nice to have) Familiarity with agentic systems and their attack surface — tool calls, retrieval, memory, multi-agent orchestration.
Expected Background: ML
- Machine learning tools: pytorch, huggingface, transformers, datasets.
- Applied deep learning and LLM experience.
- Training and evaluating deep models.
- (nice to have) finetuning LLMs, multi-modal LLMs.
- (nice to have) Familiarity with ML[NLP,LLM,Vision] interpretability methods, sparse autoencoders, linear probes — as a way of locating failure modes, not as an end in itself.
Expected Background: Software Engineering
- Development environments and tools:
- unix, git, basic clouds usage on AWS and/or GCP
- jupyter
- Programming:
- python
- (nice to have) “programming languages well-roundedness”
- experience in statically-typed and functional languages
Compensation & Benefits
- Market aligned compensation for interns in the bay area.
Requirements
- Must be authorized to work in the USA or must be able to obtain CPT (Curricular Practical Training) approval from host university.
Get AI Red Teaming Intern jobs like this→
New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.
Email me new jobsSimilar jobs
Frequently asked questions
What skills are required for AI Red Teaming Intern at realm-labs?
The required skills for AI Red Teaming Intern at realm-labs include: Machine Learning, Deep Learning, LLM, PyTorch, Hugging Face, Python, Git, AWS, GCP.
What is the seniority level for AI Red Teaming Intern at realm-labs?
AI Red Teaming Intern at realm-labs is a Internship level position.
How do I apply for AI Red Teaming Intern at realm-labs?
You can view the full description and apply for AI Red Teaming Intern at realm-labs on EchoJobs: https://echojobs.io/job/realm-labs-intern-ai-red-teaming-fall-2026-qp4vt.

